Mamba-Enhanced Vision-Language Framework for Intelligent Polyp Detection and Clinical Endoscopy Report Generation
Keywords:
polyp detection; endoscopy report generation; vision-language models; state space models; Mamba architectures; medical image segmentation; clinical decision support; deployment governanceAbstract
Colorectal cancer prevention depends heavily on the reliable identification and removal of precancerous polyps during endoscopic examination. Although deep learning has improved polyp detection and segmentation, many clinical systems remain narrowly task-oriented and do not address the interpretive documentation workload that follows a procedure. This paper presents a system-level analysis of a Mamba-enhanced vision-language framework that unifies intelligent polyp detection with structured clinical endoscopy report generation. The framework combines selective state space modeling for efficient long-range visual encoding with a vision-language alignment module and a constrained language decoder. The design is discussed from structural, infrastructural, and governance perspectives rather than as a purely algorithmic contribution. The analysis examines how linear-time sequence modeling can reduce memory burdens in long endoscopic video streams while preserving global context for polyp localization. It further considers how language generation can be coupled to visual evidence to produce preliminary reports that support rather than replace clinical judgment. Key trade-offs are addressed across detection sensitivity, false positive burden, language grounding, hallucination control, interoperability with electronic health records, and regulatory compliance. The paper also evaluates fairness risks arising from dataset composition and deployment heterogeneity, and it proposes governance mechanisms that may support safer translation into real clinical environments. The discussion emphasizes that the value of such a framework lies less in isolated benchmark performance than in its ability to operate within an accountable, auditable, and sustainable sociotechnical infrastructure.
References
1. Gu, A., & Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752.
2. Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., & Wang, X. (2024). Vision Mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417.
3. Ma, J., Li, F., & Wang, B. (2024). U-Mamba: Enhancing long-range dependency for biomedical image segmentation. arXiv preprint arXiv:2401.04722.
4. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., ... & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (pp. 8748-8763). PMLR.
5. Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., ... & Gao, J. (2024). LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36.
6. Wang, P., Xiao, X., Glissen Brown, J. R., Berzin, T. M., Tu, M., Xiong, F., ... & Liu, X. (2018). Development and validation of a deep-learning algorithm for the detection of polyps during colonoscopy. Nature Biomedical Engineering, 2(10), 741-748.
7. Jha, D., Smedsrud, P. H., Riegler, M. A., Halvorsen, P., de Lange, T., Johansen, D., & Johansen, H. D. (2020). Kvasir-SEG: A segmented polyp dataset. In Proceedings of the 26th International Conference on Multimedia Modeling (pp. 451-462). Springer.
8. Tang, Y., Chang, C., Zhao, Z., Gu, Z., & Li, Z. (2026, March). Intelligent Detection and Segmentation of Gastrointestinal Polyps Based on YOLOv11 Deep Learning. In 2026 9th International Conference on Advanced Algorithms and Control Engineering (ICAACE) (pp. 825-828). IEEE.
9. Zhang, Z., Ma, Q., Zhang, T., Chen, J., Zheng, H., & Gao, W. (2025). Switch-UMamba: Dynamic scanning vision Mamba UNet for medical image segmentation. Medical Image Analysis, 103792.
10. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (pp. 5998-6008).
11. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT (pp. 4171-4186).
12. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., ... & Lample, G. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
13. Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems (pp. 1877-1901).
14. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., ... & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730-27744.
15. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., ... & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 4015-4026).
16. Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (pp. 234-241). Springer.
17. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770-778).
18. Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. In International Conference on Learning Representations.
19. Wiens, J., Saria, S., Sendak, M., Ghassemi, M., Liu, V. X., Doshi-Velez, F., ... & Goldenberg, A. (2019). Do no harm: a roadmap for responsible machine learning for health care. Nature Medicine, 25(9), 1337-1340.
20. Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New England Journal of Medicine, 380(14), 1347-1358.
21. Gerke, S., Minssen, T., & Cohen, G. (2020). Ethical and legal challenges of artificial intelligence-driven healthcare. In Artificial Intelligence in Healthcare (pp. 295-336). Academic Press.
22. Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J., & Denniston, A. K. (2020). Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nature Medicine, 26(9), 1364-1374.
23. U.S. Food and Drug Administration. (2021). Artificial intelligence/machine learning (AI/ML)-based software as a medical device (SaMD) action plan. U.S. Department of Health and Human Services.
24. Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., ... & Dean, J. (2019). A guide to deep learning in healthcare. Nature Medicine, 25(1), 24-29.
25. Zhou, S. K., Greenspan, H., Davatzikos, C., Duncan, J. S., Van Ginneken, B., Madabhushi, A., ... & Summers, R. M. (2021). A review of deep learning in medical imaging: Imaging traits, technology trends, case studies with progress highlights, and future promises. Proceedings of the IEEE, 109(5), 820-838.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Bioinformatics Insights and Analytics

This work is licensed under a Creative Commons Attribution 4.0 International License.