Integrating Explainability, Robustness, and Regulatory Compliance in Large Language Models and Foundation AI Systems: A Comprehensive Review and Operational Framework

Bharat Joshi, Ravindra Gowda, Vikas Hegde

Abstract


ABSTRACT The rapid proliferation and industrial deployment of Large Language Models (LLMs) and Foundation AI Systems have catalyzed transformative advancements across natural language processing, decision support, autonomous reasoning, and predictive analytics. However, as these deep neural architectures increase in parameter scale and operational ubiquity, their black-box nature, vulnerability to adversarial manipulations, non deterministic behaviors, and propensity for generating hallucinated or biased outputs present critical bottlenecks to safe adoption. Addressing these multi faceted challenges requires an integrated methodology combining Explainable AI (XAI), empirical adversarial robustness, and rigorous compliance with evolving global artificial intelligence regulations (such as the European Union AI Act, US NIST AI Risk Management Framework, and ISO/IEC 42001). This paper presents a comprehensive state-of-the-art review on the technical convergence of explainability, robustness, and regulatory compliance in foundation models. We systematically evaluate feature attribution techniques, mechanistic interpretability, latent concept probing, and attention dynamics alongside safety alignment paradigms including Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimization (DPO), and runtime guardrailing. Furthermore, we analyze adversarial vulnerabilities including prompt injection, jailbreaking, vector store poisoning, and out-of-distribution performance degradation. Bridging technical implementations with legal mandates, we propose a unified architectural framework for real-time compliance auditing and continuous model monitoring. Finally, we discuss key trade-offs between model interpretability and predictive capacity, outline current technical limitations, and present high-impact directions for future research in trustworthy AI engineering.

KEYWORDS: Large Language Models, Explainable AI (XAI), Mechanistic Interpretability, Adversarial Robustness, Regulatory Compliance, EU AI Act, Safety Guardrails, Trustworthy AI.


Full Text:

PDF 92-107

Refbacks

  • There are currently no refbacks.