Comparing optimal transport and machine learning approaches for databases merging in scenarios involving missing data in covariates. Application to Medical Research

Clicks: 2
ID: 324599
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #369 of 821 articles by views in BMC Bioinformatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 821 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
MOTIVATION: One of the challenges encountered when merging heterogeneous observational medical datasets is the recoding of categorical target variables that may have been measured differently across data sources. This study compares standard machine learning-based approaches, specifically Multiple Imputation by Chained Equations (MICE), k-Nearest Neighbours (kNN), missForest, and Factor Analysis of Mixed Data (missMDA), with an Optimal Transport-based algorithm (OTrecod). Their empirical performance is evaluated in realistic data integration settings that incorporate missing covariate values, non-linear relationships, and imbalanced groups, all of which remain underexplored. RESULTS: A comprehensive simulation study was conducted, varying sample size, group imbalance, signal-to-noise ratio, non-linear constraints, and missing data mechanisms and proportions. The results reveal two distinct performance tiers, with missForest, missMDA, and OTrecod often outperforming MICE and kNN. While missForest achieves the highest recoding accuracy at low missingness levels, OTrecod and missMDA show superior robustness in high missingness scenarios. Furthermore, OTrecod excels under severe non-linear constraints. These findings are further supported by subsets of the National Child Development Study, in which OTrecod produced the most stable and consistent recoding alignments across methods. AVAILABILITY AND IMPLEMENTATION: The source code supporting this study is publicly available at https://github.com/FloAI/CompareOT and archived on Zenodo (DOI: 10.5281/zenodo.20542318).
Reference Key
openalex_W7202115518 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Flore N'kam Suguem, Sébastien Dejean, Philippe Pierre, Nicolas Savy
Journal BMC Bioinformatics
Year 2026
DOI
10.1093/bioinformatics/btag594
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.