Critical Assessment of ML models for ADMET Prediction in TDC leaderboards
Тип публікації :
Препринт
Дата випуску :
28 лютого 2026 р.
Автор(и) :
Ihor Koleiev
Roman Stratiichuk
Nazar Shevchuk
Mykola Melnychenko
Oleksiy Nyporko
Daniil Todoryshyn
Vladyslav Husak
Sergiy A. Starosyla
Semen O Yesylevskyy
Alan Nafiiev
eKNUTSHIR URL :
Журнал :
bioRxiv (Cold Spring Harbor Laboratory)
Цитування :
[APA 7] Ihor, K., Roman, S., Nazar, S., Mykola, M., Oleksiy, N., Daniil, T., Vladyslav, H., Sergiy, A. S., Semen, O. Y., & Alan, N. (2026). Critical Assessment of ML models for ADMET Prediction in TDC leaderboards. bioRxiv (Cold Spring Harbor Laboratory),. https://doi.org/10.64898/2026.02.26.708193
[ДСТУ] Critical Assessment of ML models for ADMET Prediction in TDC leaderboards / K. Ihor та ін. bioRxiv (Cold Spring Harbor Laboratory). 2026. DOI: 10.64898/2026.02.26.708193 (дата звернення: 11.09.2026).
In this work we performed a critical assessment of the benchmarking procedures used in Therapeutics Data Commons (TDC) ADMET leaderboards, focusing on reproducibility, robustness against data leakage, and signs of test-set overfitting across all 22 TDC ADMET endpoints. For each endpoint, the top 3 leaderboard models were screened with a unified protocol: execution environment reproducibility check, data leakage assessment, verification of hyperparameter optimisation practices, and final re-evaluation of results and TDC ranking. Only 3 methods (CaliciBoost, MapLight, MapLight+GNN) passed all checks and showed overall reproducible performance, whereas most of top-ranked models exhibited unavailable code, non-reproducible execution environments, runtime incompatibilities, or various methodological flaws. In particular, we identified direct or indirect data leakages in MiniMol, GradientBoost and XGBoost models. We also used our in-house models based on the Mol2Vec architecture to investigate the consequences of deliberately overfitting on the TDC test set. It is shown that deliberate or accidental tuning on the public test set may lead to significant inflation of the model metrics and leaderboard position. Our results emphasize the urgent need for better public ADMET benchmarks with the hidden test sets, strict dataset versioning and model submission with standardized inference environments.
Файл(и) :![Ескіз]()
Вантажиться...
Формат :
Adobe PDF
Розмір :
654.57 KB
Контрольна сума :
(MD5):eec597e88f3c253c717c33c53dbcb4d1
Якщо не вказано інше, ця робота розповсюджується на умовах ліцензії Attribution 4.0 International

