Multi-Omics Data Fusion and Artificial Intelligence for Phylogenetic Classification of Medicinal Plant Resources

Authors

  • Anel Melik Department of Computer Science, Binghamton University, Binghamton, NY, USA.
  • Zachary Fleming Department of Computer Science, University of Alabama at Birmingham, Birmingham, AL, USA.
  • Karan C. Menon School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, OR, USA.
  • Malcolm M. Jarvinen Department of Electrical Engineering and Computer Science, University of Kansas, Lawrence, KS, USA.

Keywords:

multi-omics data fusion; artificial intelligence; phylogenetic classification; medicinal plant resources; data governance; fairness

Abstract

Medicinal plant resources are central to pharmaceutical innovation, traditional medicine, and biodiversity conservation, yet their phylogenetic classification remains complicated by cryptic speciation, hybridization, phenotypic plasticity, and incomplete voucher-linked reference collections. This paper examines how multi-omics data fusion combined with artificial intelligence can transform phylogenetic classification from a narrow molecular barcoding task into a systems-level inference problem. It analyzes the structural trade-offs among early, intermediate, and late data fusion architectures and discusses how genomic, transcriptomic, metabolomic, and cell wall compositional features can be jointly represented for downstream classification. The discussion emphasizes system requirements such as provenance-aware data pipelines, modular model design, uncertainty quantification, validation across heterogeneous plant lineages, and continuous monitoring after deployment. The paper argues that technical accuracy alone is insufficient. Fairness, benefit-sharing, ecological sustainability, and policy coherence must be designed into the infrastructure rather than appended after model training. Drawing on cross-domain comparisons from biomedical and ecological data science, the paper provides a forward-looking perspective on resilient, governable, and ethically grounded computational infrastructure for medicinal plant phylogenetics. The conclusion highlights the need for international data governance frameworks, reproducible model registries, and participatory approaches that include Indigenous and local knowledge systems in the curation and interpretation of medicinal plant data.

References

1. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444.

2. Bersanelli, M., Mosca, E., Remondini, D., Giampieri, E., Sala, C., Castellani, G., & Milanesi, L. (2016). Methods for the integration of multi-omics data: Mathematical aspects. BMC Bioinformatics, 17(Suppl 2), 15.

3. Argelaguet, R., Velten, B., Arnol, D., Dietrich, S., Zenz, T., Marioni, J. C., Buettner, F., Huber, W., & Stegle, O. (2018). Multi-Omics Factor Analysis—a framework for unsupervised integration of multi-omics data sets. Molecular Systems Biology, 14(6), e8124.

4. CBOL Plant Working Group. (2009). A DNA barcode for land plants. Proceedings of the National Academy of Sciences, 106(31), 12794–12797.

5. Mishra, P., Kumar, A., Nagireddy, A., Mani, D. N., Shukla, A. K., Tiwari, R., & Sundaresan, V. (2016). DNA barcoding: An efficient tool to overcome authentication challenges in the herbal market. Plant Biotechnology Journal, 14(1), 8–21.

6. Hollingsworth, P. M., Forrest, L. L., Spouge, J. L., Hajibabaei, M., Ratnasingham, S., van der Bank, M., Chase, M. W., Cowan, R. S., Erickson, D. L., Fazekas, A. J., Graham, S. W., James, K. E., Kim, K.-J., Kress, W. J., Schneider, H., van AlphenStahl, J., Barrett, S. C. H., van den Berg, C., Bogarin, D., … Little, D. P. (2011). Choosing and using a plant DNA barcode. PLoS ONE, 6(5), e19254.

7. Sumner, L. W., Mendes, P., & Dixon, R. A. (2003). Plant metabolomics: Large-scale phytochemistry in the functional genomics era. Phytochemistry, 62(6), 817–836.

8. Zitnik, M., Nguyen, F., Wang, B., Leskovec, J., Goldenberg, A., & Hoffman, M. M. (2019). Machine learning for integrating data in biology and medicine: Principles, practice, and opportunities. Information Fusion, 50, 71–91.

9. Zhang, C., Rabiee, M., Sayyari, E., & Mirarab, S. (2018). ASTRAL-III: Polynomial time species tree reconstruction from partially resolved gene trees. BMC Bioinformatics, 19(Suppl 6), 153.

10. Zhao, K., Li, G., Li, C., Jiang, T., Wang, W., & Yang, Z. (2026). Identification Markers for Salvia miltiorrhiza and Its Close Relatives Based on Cell Wall Component Characteristics. Engineered Science, 40, 2122.

11. Wang, H., Cimen, E., Singh, N., & Buckler, E. (2020). Deep learning for plant genomics and crop improvement. Current Opinion in Plant Biology, 54, 34–41.

12. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 115.

13. Rieke, N., Hancox, J., Li, W., Milletarì, F., Roth, H. R., Albarqouni, S., Bakas, S., Galtier, M. N., Landman, B. A., Maier-Hein, K., Ourselin, S., Sheller, M., Summers, R. M., Trask, A., Xu, D., Baust, M., & Cardoso, M. J. (2020). The future of digital health with federated learning. npj Digital Medicine, 3, 119.

14. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., … Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018.

15. Carroll, S. R., Garba, I., Figueroa-Rodríguez, O. L., Holbrook, J., Lovett, R., Materechera, S., Parsons, M., Raseroka, K., Rodriguez-Lonebear, D., Rowe, R., Sara, R., Walker, J. D., Anderson, J., & Hudson, M. (2020). The CARE Principles for Indigenous Data Governance. Data Science Journal, 19, 43.

16. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency, 220–229.

17. Fahlgren, N., Gehan, M. A., & Baxter, I. (2015). Lights, camera, action: High-throughput plant phenotyping is ready for a close-up. Current Opinion in Plant Biology, 24, 93–99.

18. Faith, D. P. (1992). Conservation evaluation and phylogenetic diversity. Biological Conservation, 61(1), 1–10.

19. Posey, D. A., & Dutfield, G. (1996). Beyond intellectual property: Toward traditional resource rights for indigenous peoples and local communities. International Development Research Centre.

20. Gal, Y., & Ghahramani, Z. (2016). Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. Proceedings of the 33rd International Conference on Machine Learning, 1050–1059.

21. Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774.

22. Meyer, H., & Pebesma, E. (2021). Predicting into unknown space? Estimating the area of applicability of spatial prediction models. Methods in Ecology and Evolution, 12(9), 1620–1633.

Downloads

Published

2026-07-10

How to Cite

Anel Melik, Zachary Fleming, Karan C. Menon, & Malcolm M. Jarvinen. (2026). Multi-Omics Data Fusion and Artificial Intelligence for Phylogenetic Classification of Medicinal Plant Resources. Bioinformatics Insights and Analytics, 1(2). Retrieved from https://bioinfia.org/index.php/home/article/view/207