Independent Verification as a Precondition for Trustworthy AI Deployment: The Case for a Formal AI Evaluation Discipline
Keywords:
AI, AI governance, trustworthy AI, independent AI evaluation, third-party verification, AI risk management, AI assurance.Abstract
As the healthcare, financial, legal, government, and other high-impact sectors rapidly adopt artificial intelligence (AI), there has been a growing urgency for strong systems that guarantee reliability, accountability, and trust in the system. While modern principles for AI governance include transparency, risk management, fairness, and humans in the loop, the methods for assessing AI systems are still mostly decentralized, organization-specific, and relying on developer-driven assessments. The scale of the deployment of AI and the relative immaturity of independent evaluation introduce issues about the quality and authenticity of AI’s role in the safety and mission-critical domain. This paper presents a proposal to consider independent, third-party evaluation as a structural precondition for the trustworthy deployment of AI as opposed to an optional quality assurance mechanism. With a conceptual and policy-oriented approach, the study brings together insights from the field of AI governance literature, regulatory guidance, international standards, benchmark research, and practices in financial auditing, model risk management, and clinical oversight. It explores the challenges of developer self-assessment conflicts of interest, methodological differences to evaluate the problem, and a lack of external accountability and its impact on industries where AI decisions have substantial societal and economic implications. The paper also reviews the existing regulatory framework, highlighting that the framework offers detailed guidelines on the principles for managing and accountability over AI risks but lacks a prescriptive approach on how to practically assess the performance of an AI independently. It also separates performance assessment, which is based on benchmarks, from the requirements of an accepted assessment discipline, noting that an evaluator needs to be independent; assessment methods need to be reproducible; the documentation needs to be transparent; assessment criteria need to be standardized; and evidence needs to be auditable and support regulatory compliance and organizational governance. The study finds that, contrary to what some may think, there’s no problem with evaluations, but there is one problem with an objective, independent, and professionally structured evaluation discipline that produces reliable, honest evaluations in various application fields. The paper highlights the unfilled governance space, thereby advancing existing scholarly and policy debates on credible AI, and offers a theoretical basis for the reasons why independence and standardization are crucial attributes of credible AI evaluation.References
1. Gisolfi, N. (2022). Model-Centric Verification of Artificial Intelligence
(Doctoral dissertation, Carnegie Mellon University, USA).
2. Brundage, M., Avin, S., Wang, J., Belfield, H., Krueger, G., Hadfield, G.,
... & Anderljung, M. (2020). Toward trustworthy AI development:
mechanisms for supporting verifiable claims. arxiv, 2020, 2004-
07213.
3. Mazumder, P. T. (2025). AI-Driven Anti-Money Laundering Systems
for Cybersecurity Resilience in US Financial Infrastructure: A
Framework for Real-Time Threat Detection, Regulatory Compliance
and National Security. International Journal of Humanities and
Information Technology, 7(03), 90-97.
4. Zicari, R. V., Brodersen, J., Brusseau, J., Düdder, B., Eichhorn, T.,
Ivanov, T., ... & Westerlund, M. (2021). Z-Inspection®: a process to
assess trustworthy AI. IEEE Transactions on Technology and Society,
2(2), 83-97.
5. Jacovi, A., Marasović, A., Miller, T., & Goldberg, Y. (2021, March).
Formalizing trust in artificial intelligence: Prerequisites, causes and
goals of human trust in AI. In Proceedings of the 2021 ACM conference
on fairness, accountability, and transparency (pp. 624-635).
6. Mazumder, P. T. (2025). Harnessing fintech innovations for antimoney
laundering: a data-driven approach. Available at SSRN
5259084.
7. Groß, D. M. (2024). Towards trustworthy ai: Formal verification in
machine learning (Doctoral dissertation, Sl: sn).
8. Winter, P. M., Eder, S., Weissenböck, J., Schwald, C., Doms, T., Vogt,
T., ... & Nessler, B. (2021). Trusted artificial intelligence: Towards
certification of machine learning applications. arXiv preprint
arXiv:2103.16910.
9. Mazumder, P. T. (2023). Data Driven Detection of Trade Based Money
Laundering (TBML): A predictive Analytics Framework for Securing
US supply chains and Financial Integrity. SAMRIDDHI: A Journal
of Physical Sciences, Engineering and Technology, 15(04), 423-432.
10. Zicari, R. V., Amann, J., Bruneault, F., Coffee, M., Düdder, B., Hickman,
E., ... & Wurth, R. (2022). How to assess trustworthy AI in practice.
arXiv preprint arXiv:2206.09887.
11. Stet t inger, G ., Weissensteiner, P., & K hast g ir, S . (2024).
Trustworthiness assurance assessment for high-risk AI-based
systems. IEEE Access, 12, 22718-22745.
12. Díaz Rodríguez, N., Del Ser Lorente, J., Coeckelbergh, M., López
de Prado, M., Herrera Viedma, E., & Herrera Triguero, F. (2023).
Connecting the dots in trustworthy Artificial Intelligence: From AI
principles, ethics, and key requirements to responsible AI systems
and regulation. Information Fusion, 99.
13. Kowald, D., Scher, S., Pammer-Schindler, V., Müllner, P., Waxnegger,
K., Demelius, L., ... & Kopeinik, S. (2024). Establishing and evaluating
trustworthy AI: overview and research challenges. Frontiers in Big
Data, 7, 1467222.
14. Mazumder, P. T. (2025). Explainable Machine Learning Pipelines for Customer Risk Scoring in Anti-Money Laundering: A Management
and Governance Perspective. Journal of Data Analysis and Critical
Management, 1(02), 79-90.
15. Li, B., Qi, P., Liu, B., Di, S., Liu, J., Pei, J., ... & Zhou, B. (2023).
Trustworthy AI: From principles to practices. ACM Computing
Surveys, 55(9), 1-46.
16. Hohma, E., & Lütge, C. (2023). From trustworthy principles to
a trustworthy development process: The need and elements of
trusted development of AI systems. AI, 4(4), 904-925.
17. Thiebes, S., Lins, S., & Sunyaev, A. (2021). Trustworthy artificial
intelligence: S. Thiebes et al. Electronic Markets, 31(2), 447-464.
18. Thakkar, V. B. (2025). Defining Success in Probabilistic Products:
Key Performance Indicators and Lifecycle Management for
Generative AI Applications in Enterprise. International Journal of
Technology, Management and Humanities, 11(04), 66-79.
19. Suseelan, S. (2025). Big Data Analytics in Air Traffic Management:
Enhancing Aviation Safety, Operational Efficiency, and Predictive
Decision-Making. Journal of Science Technology and Social
Transformation, 1(02), 61-73.
20. Thakkar, V. B. (2025). Algorithmic Trust, Governance, and
Integration Bottlenecks of AI Tools in Legacy Project Environments.
International Journal of Humanities and Information Technology,
7(01), 80-89.
21. Suseelan, S. (2023). Explainable AI (XAI) for FAA Certifiable Aviation
Systems. SAMRIDDHI: A Journal of Physical Sciences, Engineering and
Technology, 15(04), 441-451.
22. Thakkar, V. B. (2026). From linear logistics to neural supply
chains: Predictive machine learning and the rise of autonomous
supply chain intelligence. International Journal of AI, BigData,
Computational and Management Studies, 7(1), 153-161.