Volume 13
Issue 6
IEEE/CAA Journal of Automatica Sinica
| Citation: | Q. Du, Z. Lei, J. Cheng, M. Omura, H. Hasegawa, and S. Gao, “High-order interaction and low-order parallelization of features fusion with novel Mamba-UNet architecture for medical image segmentation,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1378–1391, Jun. 2026. doi: 10.1109/JAS.2026.125720 |
| [1] |
J. E. Iglesias and M. R. Sabuncu, “Multi-atlas segmentation of biomedical images: A survey,” Med. Image Anal., vol. 24, no. 1, pp. 205–219, Aug. 2015. doi: 10.1016/j.media.2015.06.012
|
| [2] |
D.-T. Lin, C.-C. Lei, and S.-W. Hung, “Computer-aided kidney segmentation on abdominal CT images,” IEEE Trans. Inf. Technol. Biomed., vol. 10, no. 1, pp. 59–65, Jan. 2006. doi: 10.1109/TITB.2005.855561
|
| [3] |
J. Yanase and E. Triantaphyllou, “A systematic survey of computer-aided diagnosis in medicine: Past and present developments,” Exp. Syst. Appl., vol. 138, Art. no. 112821, Dec. 2019. doi: 10.1016/j.eswa.2019.112821
|
| [4] |
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. W. M. Van Der Laak, B. Van Ginneken, and C. I. Sánchez, “A survey on deep learning in medical image analysis,” Med. Image Anal., vol. 42, pp. 60–88, Dec. 2017. doi: 10.1016/j.media.2017.07.005
|
| [5] |
N. Salpea, P. Tzouveli, and D. Kollias, “Medical image segmentation: A review of modern architectures,” in Proc. European Conf. Computer Vision, Tel Aviv, Israel, 2022, pp. 691−708.
|
| [6] |
M. Antonelli, A. Reinke, S. Bakas, K. Farahani, A. Kopp-Schneider, B. A. Landman, et al., “The medical segmentation decathlon,” Nat. Commun., vol. 13, no. 1, Art. no. 4128, Jul. 2022. doi: 10.1038/s41467-022-30695-9
|
| [7] |
R. Wang, T. Lei, R. Cui, B. Zhang, H. Meng, and A. K. Nandi, “Medical image segmentation using deep learning: A survey,” IET Image Process., vol. 16, no. 5, pp. 1243–1267, Apr. 2022. doi: 10.1049/ipr2.12419
|
| [8] |
D. L. Pham, C. Xu, and J. L. Prince, “Current methods in medical image segmentation,” Annu. Rev. Biomed. Eng., vol. 2, pp. 315–337, Aug. 2000. doi: 10.1146/annurev.bioeng.2.1.315
|
| [9] |
M. H. Hesamian, W. Jia, X. He, and P. Kennedy, “Deep learning techniques for medical image segmentation: Achievements and challenges,” J. Digit. Imaging, vol. 32, no. 4, pp. 582–596, Aug. 2019. doi: 10.1007/s10278-019-00227-x
|
| [10] |
T. Dhar, N. Dey, S. Borra, and R. S. Sherratt, “Challenges of deep learning in medical image analysis—improving explainability and trust,” IEEE Trans. Technol. Soc., vol. 4, no. 1, pp. 68–75, Mar. 2023. doi: 10.1109/TTS.2023.3234203
|
| [11] |
H. Zhang, S. Cholleti, S. A. Goldman, and J. E. Fritts, “Meta-Evaluation of image segmentation using machine learning,” in Proc. IEEE Computer Society Conf. Computer Vision and Pattern Recognition, New York, USA, 2006, pp. 1138−1145.
|
| [12] |
T. A. Soomro, L. Zheng, A. J. Afifi, A. Ali, S. Soomro, M. Yin, and J. Gao, “Image segmentation for MR brain tumor detection using machine learning: A review,” IEEE Rev. Biomed. Eng., vol. 16, pp. 70−90, 2023.
|
| [13] |
S. Minaee, Y. Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image segmentation using deep learning: A survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 7, pp. 3523–3542, Jul. 2022.
|
| [14] |
Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou, “A survey of convolutional neural networks: Analysis, applications, and prospects,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 12, pp. 6999–7019, Dec. 2022.
|
| [15] |
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proc. 18th Int. Conf. Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 2015, pp. 234−241.
|
| [16] |
V. Badrinarayanan, A. Kendall, and R. Cipolla, “SegNet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 12, pp. 2481–2495, Dec. 2017. doi: 10.1109/TPAMI.2016.2644615
|
| [17] |
F. I. Diakogiannis, F. Waldner, P. Caccetta, and C. Wu, “ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data,” ISPRS J. Photogramm. Remote Sens., vol. 162, pp. 94–114, Apr. 2020. doi: 10.1016/j.isprsjprs.2020.01.013
|
| [18] |
O. Oktay, J. Schlemper, L. Le Folgoc, M. Lee, M. Heinrich, K. Misawa, et al., “Attention U-Net: Learning where to look for the pancreas,” arXiv preprint arXiv: 1804.03999, 2018.
|
| [19] |
J. Long, M. Li, and X. Wang, “Integrating spatial details with long-range contexts for semantic segmentation of very high-resolution remote-sensing images,” IEEE Geosci. Remote Sens. Lett., vol. 20, Art. no. 2501605, 2023.
|
| [20] |
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. 9th Int. Conf. Learning Representations, 2021.
|
| [21] |
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “SegFormer: Simple and efficient design for semantic segmentation with transformers,” in Proc. 35th Int. Conf. Neural Information Processing Systems, 2021, pp.12077–12090.
|
| [22] |
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, et al., “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Nashville, USA, 2021, pp. 6877−6886.
|
| [23] |
R. Strudel, R. Garcia, I. Laptev, and C. Schmid, “Segmenter: Transformer for semantic segmentation,” in Proc. IEEE/CVF Int. Conf. Computer Vision, Montreal, Canada, 2021, pp. 7242−7252.
|
| [24] |
J. Chen, J. Mei, X. Li, Y. Lu, Q. Yu, Q. Wei, et al., “TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers,” Med. Image Anal., vol. 97, Art. no. 103280, Oct. 2024. doi: 10.1016/j.media.2024.103280
|
| [25] |
L. Wang, R. Li, C. Zhang, S. Fang, C. Duan, X. Meng, and P. M. Atkinson, “UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery,” ISPRS J. Photogramm. Remote Sens., vol. 190, pp. 196–214, Aug. 2022. doi: 10.1016/j.isprsjprs.2022.06.008
|
| [26] |
H. You, Y. Xiong, X. Dai, B. Wu, P. Zhang, H. Fan, P. Vajda, and Y. C. Lin, “Castling-ViT: Compressing self-attention via switching towards linear-angular attention at vision transformer inference,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Vancouver, Canada, 2023, pp. 14431−14442.
|
| [27] |
F. Babiloni, I. Marras, J. Deng, F. Kokkinos, M. Maggioni, G. Chrysos, P. Torr, and S. Zafeiriou, “Linear complexity self-attention with 3rd order polynomials,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 11, pp. 12726–12737, Nov. 2023.
|
| [28] |
M.-H. Guo, Z.-N. Liu, T.-J. Mu, and S.-M. Hu, “Beyond self-attention: External attention using two linear layers for visual tasks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 5, pp. 5436–5447, May 2023. doi: 10.1109/tpami.2022.3211006
|
| [29] |
A. Gu, Modeling Sequences with Structured State Spaces. Stanford, USA: Stanford University, 2023.
|
| [30] |
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré, “Combining recurrent, convolutional, and continuous-time models with linear state-space layers,” in Proc. 35th Int. Conf. Neural Information Processing Systems, 2021, Art. no. 44.
|
| [31] |
A. Gu, K. Goel, and C. Ré, “Efficiently modeling long sequences with structured state spaces,” in Proc. 10th Int. Conf. Learning Representations, 2022.
|
| [32] |
K. Goel, A. Gu, C. Donahue, and C. Ré, “It’s raw! Audio generation with state-space models,” in Proc. 39th Int. Conf. Machine Learning, Baltimore, USA, 2022, pp. 7616−7633.
|
| [33] |
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv: 2312.00752, 2024.
|
| [34] |
J. Ma, F. Li, and B. Wang, “U-Mamba: Enhancing long-range dependency for biomedical image segmentation,” arXiv preprint arXiv: 2401.04722, 2024.
|
| [35] |
Z. Wang, J.-Q. Zheng, Y. Zhang, G. Cui, and L. Li, “Mamba-UNet: Unet-like pure visual Mamba for medical image segmentation,” arXiv preprint arXiv: 2402.05079, 2024.
|
| [36] |
C. Jiang, R. Wu, Y. Liu, Y. Wang, Q. Chang, P. Liang, and Y. Fan, “A high-order focus interaction model and oral ulcer dataset for oral ulcer segmentation,” Sci. Rep., vol. 14, no. 1, Art. no. 20085, Aug. 2024. doi: 10.1038/s41598-024-69125-9
|
| [37] |
R. Wu, Y. Liu, P. Liang, and Q. Chang, “Only Positive Cases: 5-fold high-order attention interaction model for skin segmentation derived classification,” arXiv preprint arXiv: 2311.15625, 2023.
|
| [38] |
Y. Rao, W. Zhao, Y. Tang, J. Zhou, S. N. Lim, and J. Lu, “HorNet: Efficient high-order spatial interactions with recursive gated convolutions,” in Proc. 36th Int. Conf. Neural Information Processing Systems, New Orleans, USA, 2022, pp. 10353−10366.
|
| [39] |
X. Liu, C. Zhang, and L. Zhang, “Vision Mamba: A comprehensive survey and taxonomy,” arXiv preprint arXiv: 2405.04404, 2024.
|
| [40] |
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, Boston, USA, 2015, pp. 3431−3440.
|
| [41] |
Y. Ye, P. Huang, Y. Sun, and D. Shi, “MBSNet: A deep learning model for multibody dynamics simulation and its application to a vehicle-track system,” Mech. Syst. Signal Process., vol. 157, Art. no. 107716, Aug. 2021. doi: 10.1016/j.ymssp.2021.107716
|
| [42] |
R. Wu, P. Liang, X. Huang, L. Shi, Y. Gu, H. Zhu, and Q. Chang, “MHorUNet: High-order spatial interaction UNet for skin lesion segmentation,” Biomed. Signal Process. Control, vol. 88, Art. no. 105517, Feb. 2024. doi: 10.1016/j.bspc.2023.105517
|
| [43] |
R. Wu, H. Lv, P. Liang, X. Cui, Q. Chang, and X. Huang, “HSH-UNet: Hybrid selective high order interactive U-shaped model for automated skin lesion segmentation,” Comput. Biol. Med., vol. 168, Art. no. 107798, Jan. 2024. doi: 10.1016/j.compbiomed.2023.107798
|
| [44] |
S. Mehta, M. Rastegari, A. Caspi, L. Shapiro, and H. Hajishirzi, “ESPNet: Efficient spatial pyramid of dilated convolutions for semantic segmentation,” in Proc. 15th European Conf. Computer Vision, Munich, Germany, 2018, pp. 561−580.
|
| [45] |
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. 31st Int. Conf. Neural Information Processing Systems, Long Beach, USA, 2017, pp. 6000−6010.
|
| [46] |
H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin-Unet: Unet-like pure transformer for medical image segmentation,” in Proc. European Conf. Computer Vision, Tel Aviv, Israel, 2022, pp. 205−218.
|
| [47] |
Y. Zhang, H. Liu, and Q. Hu, “TransFuse: Fusing transformers and CNNs for medical image segmentation,” in Proc. 24th Int. Conf. Medical Image Computing and Computer-Assisted Intervention, Strasbourg, France, 2021, pp. 14−24.
|
| [48] |
W. Zhou, H. Wu, and Q. Jiang, “MDNet: Mamba-effective diffusion-distillation network for RGB-thermal urban dense prediction,” IEEE Trans. Circuits Syst. Video Technol., vol. 35, no. 4, pp. 3222–3233, Apr. 2025. doi: 10.1109/TCSVT.2024.3508058
|
| [49] |
M. Ju, S. Xie, and F. Li, “Improving skip connection in U-Net through fusion perspective with Mamba for image dehazing,” IEEE Trans. Consum. Electron., vol. 70, no. 4, pp. 7505–7514, Nov. 2024. doi: 10.1109/TCE.2024.3417476
|
| [50] |
Q. Liu, J. Yue, Y. Fang, S. Xia, and L. Fang, “HyperMamba: A spectral-spatial adaptive Mamba for hyperspectral image classification,” IEEE Trans. Geosci. Remote Sens., vol. 62, Art. no. 5536514, 2024.
|
| [51] |
J. Ruan, J. Li, and S. Xiang, “VM-UNet: Vision Mamba UNet for medical image segmentation,” ACM Trans. Multimedia Comput. Commun. Appl., 2025, DOI: 10.1145/3767748.
|
| [52] |
R. Wu, Y. Liu, G. Ning, P. Liang, and Q. Chang, “Ultralight VM-UNet: Parallel vision Mamba significantly reduces parameters for skin lesion segmentation,” Patterns, vol. 6, no. 11, Art. no. 101298, Nov. 2025. doi: 10.1016/j.patter.2025.101298
|
| [53] |
Z. Xing, T. Ye, Y. Yang, G. Liu, and L. Zhu, “SegMamba: Long-range sequential modeling Mamba for 3D medical image segmentation,” in Proc. 27th Int. Conf. Medical Image Computing and Computer-Assisted Intervention, Marrakesh, Morocco, 2024, pp. 578−588.
|
| [54] |
C. Zheng, J. Nie, Z. Wang, N. Song, J. Wang, and Z. Wei, “High-order semantic decoupling network for remote sensing image semantic segmentation,” IEEE Trans. Geosci. Remote Sens., vol. 61, Art. no. 5401415, Feb. 2023.
|
| [55] |
X. Sun, Y. Zhang, C. Chen, S. Xie, and J. Dong, “High-order paired-ASPP for deep semantic segmentation networks,” Inf. Sci., vol. 646, Art. no. 119364, Oct. 2023. doi: 10.1016/j.ins.2023.119364
|
| [56] |
K. Zhang, Y. Wu, M. Dong, B. Liu, D. Liu, and Q. Liu, “Deep object co-segmentation and co-saliency detection via high-order spatial-semantic network modulation,” IEEE Trans. Multimed., vol. 25, pp. 5733−5746, 2023.
|
| [57] |
R. Wu, Y. Liu, P. Liang, and Q. Chang, “H-vmunet: High-order vision Mamba UNet for medical image segmentation,” Neurocomputing, vol. 624, Art. no. 129447, Apr. 2025. doi: 10.1016/j.neucom.2025.129447
|
| [58] |
J. Ruan, S. Xiang, M. Xie, T. Liu, and Y. Fu, “MALUNet: A multi-attention and light-weight UNet for skin lesion segmentation,” in Proc. IEEE Int. Conf. Bioinformatics and Biomedicine, Las Vegas, USA, 2022, pp. 1150−1156.
|
| [59] |
S. Nishida, T. Ledgeway, and M. Edwards, “Dual multiple-scale processing for motion in the human visual system,” Vis. Res., vol. 37, no. 19, pp. 2685–2698, Oct. 1997. doi: 10.1016/S0042-6989(97)00092-8
|
| [60] |
Z. Liu, P. Ma, D. Chen, W. Pei, and Q. Ma, “Scale-teaching: Robust multi-scale training for time series classification with noisy labels,” Proc. 37th Int. Conf. Neural Information Processing System, New Orleans, USA, 2023, pp.33726−33757.
|
| [61] |
M. Rahman, M. Munir, and R. Marculescu, “EMCAD: Efficient multi-scale convolutional attention decoding for medical image segmentation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2024, pp. 11769−11779.
|
| [62] |
I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, et al., “MLP-Mixer: An all-MLP architecture for vision,” in Proc. 35th Int. Conf. Neural Information Processing System, 2021, Art. no. 1857.
|
| [63] |
N. C. F. Codella, D. Gutman, M. E. Celebi, B. Helba, M. A. Marchetti, S. W. Dusza, et al., “Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (ISIC),” in Proc. IEEE 15th Int. Symp. Biomedical Imaging, Washington, USA, 2018, pp. 168−172.
|
| [64] |
N. Codella, V. Rotemberg, P. Tschandl, M. E. Celebi, S. Dusza, D. Gutman, et al., “Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (ISIC),” arXiv preprint arXiv: 1902.03368, 2019.
|
| [65] |
T. Mendonça, P. M. Ferreira, J. S. Marques, A. R. S. Marcal, and J. Rozeira, “PH2-A dermoscopic image database for research and benchmarking,” in Proc. 35th Annu. Int. Conf. IEEE Engineering in Medicine and Biology Society, Osaka, Japan, 2013, pp. 5437−5440.
|
| [66] |
Z. Zhuang, N. Li, A. N. Joseph Raj, V. G. V. Mahesh, and S. Qiu, “An RDAU-NET model for lesion segmentation in breast ultrasound images,” PLoS One, vol. 14, no. 8, Art. no. e0221535, 2019. doi: 10.1371/journal.pone.0221535
|
| [67] |
P. Zhou, X. Xie, Z. Lin, and S. Yan, “Towards understanding convergence and generalization of AdamW,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 9, pp. 6486–6493, Sep. 2024. doi: 10.1109/TPAMI.2024.3382294
|
| [68] |
H. Wu, Z. Zhao, and Z. Wang, “META-Unet: Multi-scale efficient transformer attention Unet for fast and high-accuracy polyp segmentation,” IEEE Trans. Autom. Sci. Eng., vol. 21, no. 3, pp. 4117–4128, Jul. 2024. doi: 10.1109/TASE.2023.3292373
|
| [69] |
A. Kumar, Y. Guo, X. Huang, L. Ren, and X. Liu, “SeaBird: Segmentation in bird’s view with dice loss improves monocular 3D detection of large objects,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2024, pp. 10269−10280.
|
| [70] |
T. Eelbode, J. Bertels, M. Berman, D. Vandermeulen, F. Maes, R. Bisschops, and M. B. Blaschko, “Optimization for medical image segmentation: Theory and practice when evaluating with dice score or jaccard index,” IEEE Trans. Med. Imaging, vol. 39, no. 11, pp. 3679–3690, Nov. 2020. doi: 10.1109/TMI.2020.3002417
|
| [71] |
S. Zou, M. Zhang, B. Fan, Z. Zhou, and X. Zou, “SkinMamba: A precision skin lesion segmentation architecture with cross-scale global state modeling and frequency boundary guidance,” arXiv preprint arXiv: 2409.10890, 2024.
|
| [72] |
J.-H. Nam, N. S. Syazwany, S. J. Kim, and S.-C. Lee, “Modality-agnostic domain generalizable medical image segmentation by Multi-Frequency in multi-scale attention,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2024, pp. 11480−11491.
|
| [73] |
X. Shen, J. Yang, C. Wei, B. Deng, J. Huang, X.-S. Hua, X. Cheng, and K. Liang, “DCT-Mask: Discrete cosine transform mask representation for instance segmentation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Nashville, USA, 2021, pp. 8716−8725.
|
| [74] |
D. Ravì, M. Bober, G. M. Farinella, M. Guarnera, and S. Battiato, “Semantic segmentation of images exploiting DCT based features and random forest,” Pattern Recogn., vol. 52, pp. 260–273, Apr. 2016. doi: 10.1016/j.patcog.2015.10.021
|
JAS-2025-0834-Supp.pdf
|
|