Best Practices in AI and Data Science Models Evaluation
Abstract: Evaluating Artificial Intelligence (AI) and data science models is crucial to ensure their reliability, fairness, and applicability in real-world scenarios. This paper highlights best practices for model evaluation, emphasizing the importance of selecting appropriate metrics aligned with business or research goals. Key considerations include using robust validation strategies (e.g., cross-validation), monitoring for overfitting, and ensuring data splits preserve class distributions. Fairness, interpretability, and reproducibility are essential, particularly in high-stakes domains like healthcare or finance. Additionally, evaluating models across multiple datasets or demographic subgroups helps uncover biases and improve generalizability. Adopting standardized reporting practices and open-source benchmarks further strengthens the evaluation process. By adhering to these practices, practitioners can build more trustworthy and effective AI systems
Show BibTeX
@inproceedings{DBLP:conf/gi/BanerjeeTU25,
author = {Debayan Banerjee and
Tilahun Abedissa Taffa and
Ricardo Usbeck},
editor = {Ulrike Lucke and
Stefan Stieglitz and
Falk Uebernickel and
Anna{-}Lena Lamprecht and
Maike Klein},
title = {Best Practices in {AI} and Data Science Models Evaluation},
booktitle = {55. Jahrestagung der Gesellschaft f{\"{u}}r Informatik, {INFORMATIK}
2025: The Wide Open - Offenheit von Source bis Science, Potsdam, Germany,
September 16-19, 2025},
series = {{LNI}},
volume = {{P-366}},
pages = {1211--1219},
publisher = {Gesellschaft f{\"{u}}r Informatik, Bonn},
year = {2025},
url = {https://doi.org/10.18420/inf2025\_105},
doi = {10.18420/INF2025\_105},
timestamp = {Wed, 29 Oct 2025 15:57:22 +0100},
biburl = {https://dblp.org/rec/conf/gi/BanerjeeTU25.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}