Independent Verification as a Precondition for Trustworthy AI Deployment: The Case for a Formal AI Evaluation Discipline

Authors

  • Isaac Kubvoruno Mindrac AI Author

Keywords:

AI, AI governance, trustworthy AI, independent AI evaluation, third-party verification, AI risk management, AI assurance.

Abstract

As the healthcare, financial, legal, government, and other high-impact sectors rapidly adopt artificial intelligence (AI), there has been a growing urgency for strong systems that guarantee reliability, accountability, and trust in the system. While modern principles for AI governance include transparency, risk management, fairness, and humans in the loop, the methods for assessing AI systems are still mostly decentralized, organization-specific, and relying on developer-driven assessments. The scale of the deployment of AI and the relative immaturity of independent evaluation introduce issues about the quality and authenticity of AI’s role in the safety and mission-critical domain. This paper presents a proposal to consider independent, third-party evaluation as a structural precondition for the trustworthy deployment of AI as opposed to an optional quality assurance mechanism. With a conceptual and policy-oriented approach, the study brings together insights from the field of AI governance literature, regulatory guidance, international standards, benchmark research, and practices in financial auditing, model risk management, and clinical oversight. It explores the challenges of developer self-assessment conflicts of interest, methodological differences to evaluate the problem, and a lack of external accountability and its impact on industries where AI decisions have substantial societal and economic implications. The paper also reviews the existing regulatory framework, highlighting that the framework offers detailed guidelines on the principles for managing and accountability over AI risks but lacks a prescriptive approach on how to practically assess the performance of an AI independently. It also separates performance assessment, which is based on benchmarks, from the requirements of an accepted assessment discipline, noting that an evaluator needs to be independent; assessment methods need to be reproducible; the documentation needs to be transparent; assessment criteria need to be standardized; and evidence needs to be auditable and support regulatory compliance and organizational governance. The study finds that, contrary to what some may think, there’s no problem with evaluations, but there is one problem with an objective, independent, and professionally structured evaluation discipline that produces reliable, honest evaluations in various application fields. The paper highlights the unfilled governance space, thereby advancing existing scholarly and policy debates on credible AI, and offers a theoretical basis for the reasons why independence and standardization are crucial attributes of credible AI evaluation.

References

1. Gisolfi, N. (2022). Model-Centric Verification of Artificial Intelligence

(Doctoral dissertation, Carnegie Mellon University, USA).

2. Brundage, M., Avin, S., Wang, J., Belfield, H., Krueger, G., Hadfield, G.,

... & Anderljung, M. (2020). Toward trustworthy AI development:

mechanisms for supporting verifiable claims. arxiv, 2020, 2004-

07213.

3. Mazumder, P. T. (2025). AI-Driven Anti-Money Laundering Systems

for Cybersecurity Resilience in US Financial Infrastructure: A

Framework for Real-Time Threat Detection, Regulatory Compliance

and National Security. International Journal of Humanities and

Information Technology, 7(03), 90-97.

4. Zicari, R. V., Brodersen, J., Brusseau, J., Düdder, B., Eichhorn, T.,

Ivanov, T., ... & Westerlund, M. (2021). Z-Inspection®: a process to

assess trustworthy AI. IEEE Transactions on Technology and Society,

2(2), 83-97.

5. Jacovi, A., Marasović, A., Miller, T., & Goldberg, Y. (2021, March).

Formalizing trust in artificial intelligence: Prerequisites, causes and

goals of human trust in AI. In Proceedings of the 2021 ACM conference

on fairness, accountability, and transparency (pp. 624-635).

6. Mazumder, P. T. (2025). Harnessing fintech innovations for antimoney

laundering: a data-driven approach. Available at SSRN

5259084.

7. Groß, D. M. (2024). Towards trustworthy ai: Formal verification in

machine learning (Doctoral dissertation, Sl: sn).

8. Winter, P. M., Eder, S., Weissenböck, J., Schwald, C., Doms, T., Vogt,

T., ... & Nessler, B. (2021). Trusted artificial intelligence: Towards

certification of machine learning applications. arXiv preprint

arXiv:2103.16910.

9. Mazumder, P. T. (2023). Data Driven Detection of Trade Based Money

Laundering (TBML): A predictive Analytics Framework for Securing

US supply chains and Financial Integrity. SAMRIDDHI: A Journal

of Physical Sciences, Engineering and Technology, 15(04), 423-432.

10. Zicari, R. V., Amann, J., Bruneault, F., Coffee, M., Düdder, B., Hickman,

E., ... & Wurth, R. (2022). How to assess trustworthy AI in practice.

arXiv preprint arXiv:2206.09887.

11. Stet t inger, G ., Weissensteiner, P., & K hast g ir, S . (2024).

Trustworthiness assurance assessment for high-risk AI-based

systems. IEEE Access, 12, 22718-22745.

12. Díaz Rodríguez, N., Del Ser Lorente, J., Coeckelbergh, M., López

de Prado, M., Herrera Viedma, E., & Herrera Triguero, F. (2023).

Connecting the dots in trustworthy Artificial Intelligence: From AI

principles, ethics, and key requirements to responsible AI systems

and regulation. Information Fusion, 99.

13. Kowald, D., Scher, S., Pammer-Schindler, V., Müllner, P., Waxnegger,

K., Demelius, L., ... & Kopeinik, S. (2024). Establishing and evaluating

trustworthy AI: overview and research challenges. Frontiers in Big

Data, 7, 1467222.

14. Mazumder, P. T. (2025). Explainable Machine Learning Pipelines for Customer Risk Scoring in Anti-Money Laundering: A Management

and Governance Perspective. Journal of Data Analysis and Critical

Management, 1(02), 79-90.

15. Li, B., Qi, P., Liu, B., Di, S., Liu, J., Pei, J., ... & Zhou, B. (2023).

Trustworthy AI: From principles to practices. ACM Computing

Surveys, 55(9), 1-46.

16. Hohma, E., & Lütge, C. (2023). From trustworthy principles to

a trustworthy development process: The need and elements of

trusted development of AI systems. AI, 4(4), 904-925.

17. Thiebes, S., Lins, S., & Sunyaev, A. (2021). Trustworthy artificial

intelligence: S. Thiebes et al. Electronic Markets, 31(2), 447-464.

18. Thakkar, V. B. (2025). Defining Success in Probabilistic Products:

Key Performance Indicators and Lifecycle Management for

Generative AI Applications in Enterprise. International Journal of

Technology, Management and Humanities, 11(04), 66-79.

19. Suseelan, S. (2025). Big Data Analytics in Air Traffic Management:

Enhancing Aviation Safety, Operational Efficiency, and Predictive

Decision-Making. Journal of Science Technology and Social

Transformation, 1(02), 61-73.

20. Thakkar, V. B. (2025). Algorithmic Trust, Governance, and

Integration Bottlenecks of AI Tools in Legacy Project Environments.

International Journal of Humanities and Information Technology,

7(01), 80-89.

21. Suseelan, S. (2023). Explainable AI (XAI) for FAA Certifiable Aviation

Systems. SAMRIDDHI: A Journal of Physical Sciences, Engineering and

Technology, 15(04), 441-451.

22. Thakkar, V. B. (2026). From linear logistics to neural supply

chains: Predictive machine learning and the rise of autonomous

supply chain intelligence. International Journal of AI, BigData,

Computational and Management Studies, 7(1), 153-161.

Published

2025-12-31

How to Cite

Independent Verification as a Precondition for Trustworthy AI Deployment: The Case for a Formal AI Evaluation Discipline. (2025). Journal of Integrated Science, Technology and Management, 1(01), 15-30. https://jistm.info/index.php/jistm/article/view/33