A defensible AI security architecture treats the entire system as a compound attack surface, not just the model endpoint. It needs authenticated policy, canonical state, and fail-closed grounding before generation, because an agent can be compromised through credentials, tools, retrieval data, dependencies, or memory even when the model itself remains available.
The production scenario is familiar. A customer-support agent can read a tenant's documents, call a ticketing API, update account records, and retain memory across conversations. A user submits an apparently ordinary request, but a retrieved document contains an instruction to ignore policy and expose another customer's data. The model produces a plausible response, the tool broker accepts a broadly scoped token, and monitoring records the event only after the unauthorized action has happened.
That failure isn't primarily a prompt problem. It's an authority and evidence problem. The model proposed an action, but no independent policy layer checked whether the principal, tenant, current state, evidence, and requested capability justified it.
Why AI Security Incidents Reveal an Architecture Gap
A production team ships a support agent with access to billing records, a ticketing API, and long-lived service credentials. Nothing looks broken in staging. Then a poisoned document lands in the retrieval index, the agent cites it as if it were trusted context, and a broadly scoped token lets a routine tool call cross a tenant boundary. The model did not need to be compromised. The authority model was already too loose.
That is why incident review keeps pointing past prompts and toward architecture. Security teams often start with the visible layer. They test jailbreaks, add prompt filters, compare model providers, and inspect outputs. Those controls still matter. But many production failures arrive through ordinary application paths such as permissive service accounts, weak integration checks, poisoned documents, or operator workflows that give an agent more authority than the task requires.
A 2025 peer-reviewed AAAI study documented 32 real-world AI security incidents and linked many of them to non-compliance with established security and privacy practices. The authors also warned that the published sample does not represent all global incidents, because public reporting often omits attacker motives, system details, and application context. That reporting gap is itself a measurement problem: teams cannot compare exposure reliably without standardized telemetry, provenance, incident taxonomies, and disclosure practices. The NIST AI Risk Management Framework is useful here because it treats governance, measurement, and operational controls as a lifecycle rather than a one-time review.
The pattern is similar in the MIT incident data. The MIT AI Incident Tracker reported that incidents with high national-security impact increased from 10 in 2023 to 29 in 2024, an increase of 19 incidents, or 190 percent. It also found that most incidents were tied to human-controlled systems rather than fully autonomous ones. The MIT AI Incident Tracker data supports a practical conclusion. Security architecture has to constrain human-operated workflows, service identities, integrations, and deployment permissions, because non-human actors inherit whatever authority people and systems hand them.
What the incident pattern changes
For engineers, the change is concrete. A malicious RAG document does not need to alter model weights. It only needs to enter ingestion, survive validation, and reach the next context window with enough authority attached to its contents. A stolen API token does not need to persuade the model either. It can call the tool directly unless the broker authenticates each request, checks scope, and records which machine identity exercised which permission.
Practical rule: Treat every model proposal as untrusted until an authenticated policy engine validates the action against current state and authorized evidence.
Public incident databases still undercount this class of failure. Many teams cannot reconstruct which document was retrieved, which identity approved the tool call, or which model version produced the response. That is an evidence problem as much as a security problem. The starting point for AI security architecture is authority, system boundaries, and proof of why an action was allowed.
Mapping the Compound Attack Surface in AI Systems
NIST's adversarial-machine-learning work places security concerns across training and test data, model behavior, model outputs, software dependencies, hardware, and surrounding network and storage systems. That taxonomy rules out an endpoint-only design. A model-serving API can be hardened while the model registry, vector database, tool broker, or memory store remains exposed.
MITRE ATLAS offers a complementary threat model for attacks against AI-enabled systems, abuse or manipulation of AI capabilities, and harmful autonomous behavior. Its published data includes 16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations, and 73 case studies. The MITRE ATLAS release data is useful because it turns a broad threat narrative into a control matrix.
A practical boundary diagram reads from left to right:
Authenticated caller → policy decision → tenant-scoped retrieval → provenance-carrying evidence → model runtime → tool broker → external system, with telemetry and incident response crossing every layer. Prompts, retrieved documents, embeddings, model files, tool arguments, and persistent memory each need their own validation and authorization decision.
Attack Surface Layers in AI Systems
| Layer | Primary Risk | Key Control |
|---|---|---|
| Training and test data | Poisoning, leakage, unauthorized modification | Dataset manifests, integrity checks, access controls, validation |
| Model artifacts | Tampering, backdoors, untrusted downloads | Registry controls, cryptographic hashes, scan-before-deploy gates |
| Runtime and dependencies | Vulnerable inference software or compromised packages | Isolation, signed releases, reproducible build metadata |
| Retrieval and memory | Cross-tenant disclosure, poisoned evidence, stale context | Execution-time authorization, provenance, freshness checks |
| Tools and integrations | Excessive authority, malicious parameters, unsafe side effects | Allowlisted broker, scoped credentials, approval gates |
| Network and storage | Exfiltration, unauthorized access, availability attacks | Segmentation, encryption, egress policy, monitoring |
| Outputs | Confabulation, sensitive disclosure, unsupported action | Claim validation, output classification, abstention |
Provenance must operate as a control, not as documentation. Record dataset and model URLs, cryptographic hashes, and, where available, PKI certificates. Maintain immutable artifact manifests, separated development and production credentials, signed releases, reproducible build metadata, and rollback-capable deployment. If an attacker alters a corpus, dependency, model file, or retrieval source without detection, the output may remain syntactically valid while becoming systematically compromised. AletheionAGI's overview of large language model security is relevant here because it frames the model as one part of a larger security boundary.
Access Control and Machine Identity for AI Agents
The hardest operational question in an agent deployment isn't “Which model should we use?” It's “Which principal is authorized to perform this action right now?” An agent that answers support questions should not automatically inherit the permissions needed to refund an order, alter a customer record, or export a data set.
Recent industry research reports that machine identities outnumber human identities by more than 80 to 1, while nearly half reportedly hold sensitive or privileged access. The same research reports that only 48% of organizations have identity governance for AI entities, and 55% have access controls for AI agents. Infosys' analysis of enterprise AI identity and access architecture highlights the gap between the number of non-human principals and the governance teams apply to them.

Build authority around the task
Use a separate identity model for each meaningful capability. An agent may receive a short-lived, per-call credential for reading an authorized knowledge source, then a different credential for proposing a ticket update. The model should never mint or select its own permissions, and a user's broad identity shouldn't be copied into an autonomous workflow.
A sound request path looks like this:
- Authenticate the human or upstream service. Establish the requesting principal, tenant, session, and purpose.
- Resolve canonical state. Load the current order, account, ticket, or workflow state from an authoritative system.
- Parse and validate intent. Convert natural language into an action with explicit target, scope, constraints, and unresolved fields.
- Issue narrow authority. Grant only the capability required for this call, with resource and time boundaries.
- Require approval for irreversible actions. The agent may propose a refund or deletion, but an authorized principal or workflow must approve it.
- Record the authorization chain. Log principal, policy decision, resource, arguments, evidence, and final result.
The trade-off is administrative complexity. Per-call credentials and policy evaluation add latency and require careful failure handling, but broad persistent credentials create a much larger blast radius. Monitoring can't repair an authorization model that permits every proposed action.
Teams responsible for identity operations may also benefit from this guide from on threat detection, particularly when connecting agent activity to existing detection and response processes. The important design boundary remains the same: a model can suggest, but an authenticated policy layer decides.
Fail-Closed Grounding and Evidence Provenance
NIST's Generative AI Profile classifies confabulation as AI-generated false or misleading content, including what is commonly called hallucination. In a security-sensitive application, that makes grounding more than an accuracy feature. A generated claim that lacks authorized evidence can trigger an incorrect payment, expose a private record, or cause an operator to trust a fabricated status.
The request path should separate retrieval from generation:
Authenticated identity → policy decision → tenant-scoped retrieval → provenance-carrying evidence bundle → claim validator → response or abstention.

A provenance record should contain the tenant or data-domain identifier, source object and version, ingestion time, access policy, user or service principal, retrieval query, and exact evidence span supplied to the model. This lets an investigator distinguish “the model mentioned a source” from “the source was authorized, current, and actually entailed the claim.” NIST's Generative AI Profile discussion of data provenance supports treating provenance as a structured security object.
Where fail-closed behavior matters
An empty retrieval result must not turn into an unrestricted generation request on its own. A stale document must not outrank a current system of record. A document that fails integrity or tenant checks must not reach the model just because it contains relevant keywords.
The validator should map each material claim to an evidence span, test for contradiction and completeness, and abstain when the required support is missing. Abstention can be a clarification request, a refusal to take action, or an escalation to an operator. It isn't a model failure when the system lacks authorized evidence. It's the correct security result.
This explanation of grounding is useful for separating evidence delivery from generation. AletheionAGI's GQueries can operate as a grounding and evidence-control layer with existing LLMs, RAG pipelines, vector databases, and memory systems. The architecture still needs independent access enforcement, retention controls, source validation, and testing for stale, poisoned, or cross-tenant records.
Evaluating Retrieval Faithfulness Across Reasoning Depth
Retrieval can improve the context available to a model without guaranteeing that the final answer is faithful to that context. A systematic benchmark discussion defines faithfulness hallucinations as outputs unsupported by retrieved evidence and evaluates systems on HotpotQA and HaluBench, covering multi-hop reasoning and single-hop hallucination detection. These benchmarks are useful test environments, not proof of security for a particular customer corpus, vector store, policy engine, or memory implementation. The benchmark discussion also reports that retrieval behavior varies by task.
Retrieval Alignment by Evaluation Domain
| Setting Type | Examples | Alignment Effect |
|---|---|---|
| Knowledge-grounded settings | RAGTruth, PubMedQA | Hybrid retrieval can improve alignment |
| Single-hop hallucination detection | HaluBench | Tests whether claims are supported in a relatively direct context |
| Multi-hop reasoning | HotpotQA | Tests whether evidence can be combined across multiple steps |
| Reasoning-intensive settings | DROP, FinanceBench, CovidQA | Hybrid retrieval can increase hallucination rates |
The trade-off is easy to miss. Hybrid retrieval can improve alignment in knowledge-grounded settings such as RAGTruth and PubMedQA, yet increase hallucination rates in reasoning-intensive settings including DROP, FinanceBench, and CovidQA. A single “RAG accuracy” number hides whether the system retrieved the right evidence, used it faithfully, or crossed a permission boundary.
Measure the controls separately
Track retrieval recall, authorization correctness, evidence-to-claim entailment, citation precision, abstention behavior, and task success as distinct measures. Stratify results by single-hop and multi-hop questions, data domain, source freshness, tenant, and system version.
A production test should retrieve candidates, discard records that fail identity, tenant, freshness, or policy checks, require material claims to map to evidence spans, run contradiction checks, and return abstention or clarification when support is incomplete. A practical RAG evaluation framework should define the denominator, dataset split, model version, retrieval version, threshold, and known failure modes before a team claims improvement.
Why Authority-First Architecture Beats Model Monitoring
Monitoring is valuable, but it is a detective control. It can identify unusual tool calls, repeated refusals, unexpected retrieval patterns, or anomalous output classifications after a system has already made a decision. If an agent retains broad, persistent credentials, more monitoring may not materially reduce risk. The policy model still allows the action.
Authority-first design changes the order of operations. The system authenticates the principal, checks canonical state, validates intent, constrains the tool capability, and evaluates evidence before execution. Monitoring then records whether the authorized path behaved as expected.
A useful monitor tells you that an agent acted strangely. An authority boundary prevents the agent from acting outside its mandate.
This distinction applies across SaaS support, marketplaces, finance, and enterprise automation. A support agent may draft a response but not send it without a customer-record check. A marketplace agent may recommend a refund but not change the ledger without transaction-state validation. A finance workflow may prepare a payment request but require a separate approval principal for execution.
The design does introduce friction. Users may see more clarifying questions, operators may approve high-consequence actions, and engineering teams must maintain policy definitions as workflows change. Those costs are preferable to asking a probabilistic component to enforce authority through instructions alone. This practical discussion of zero trust principles provides useful context for authenticating and authorizing each request rather than trusting network location or prior interaction.
Threat Modeling with MITRE ATLAS and Operational Controls
An agent pulls a document from the wrong tenant, passes a hidden instruction to a tool, and opens a support refund flow. If the exercise only tests whether the model resists the instruction, the team misses the authority failure that made the action possible.
Use the ATLAS taxonomy introduced earlier to build an auditable matrix, not a narrative checklist. The point is to map each relevant technique to the control that should stop it, the signal that proves the control fired, and the playbook that contains the blast radius when it did not.
For every in-scope technique, assign:
- A control owner: The team responsible for implementation and maintenance.
- A preventive control: For example, tenant filtering, signed artifacts, scoped credentials, or an egress rule.
- A detective signal: Such as an unexpected tool argument, retrieval from another namespace, or model-hash mismatch.
- A response playbook: The action required to revoke access, quarantine artifacts, pause workflows, or restore a known-good version.
- A frozen regression test: A repeatable test that runs against a defined system and dataset version.
Take prompt injection as one example. The owner may be the platform team. The preventive control is not "better prompting." It is a broker that strips untrusted instructions from retrieved content, enforces per-call identity and scope, and blocks tool use unless policy and canonical state allow it. The signal is a mismatch between requested action and authorized capability, or a retrieval path that crossed tenant boundaries. The playbook revokes the agent token, quarantines the source corpus, and reruns the frozen test before service resumes.
Tool calls should pass through an allowlisted broker. Retrieval should authenticate sources, preserve provenance, constrain context to the requesting tenant, and detect poisoned or conflicting content.
Coverage is the operational objective. Every relevant technique should have an owner, control, signal, playbook, and repeatable test.
Log policy decisions, tool arguments, retrieved identifiers, output classifications, refusal or abstention states, and principal information without storing unnecessary sensitive payloads. A lower attack rate may only mean weak test coverage. A complete control matrix gives security teams something they can inspect, challenge, and improve.
Building a Practical Adoption Sequence for AI Security
A practical rollout starts where failures are hardest to contain. Give every agent, broker, retrieval service, and automation a distinct identity before expanding what it can do. In production, broad shared credentials hide who acted, make policy exceptions sticky, and turn a single prompt or tool-path failure into a wider access problem.
Establish the authority boundary first
Inventory agents, services, tools, tenants, data domains, and irreversible actions. Then bind each non-human actor to scoped, per-call credentials and require a separate policy decision for every retrieval, memory read, tool call, and write. This is the control point that keeps an agent from turning retrieved text into authority it never had.
Make evidence traceable
Traceability has to survive incident review. Add tenant-scoped retrieval, source versioning, ingestion timestamps, access policy, exact evidence spans, and integrity checks. Treat retrieved text as untrusted input. A prompt-injected document may suggest an action, but authorization still has to come from canonical state and authenticated policy.
Enforce unsupported-claim handling
Once authority and provenance are stable, measure whether the system can prove what it says. Implement claim-level support checks, contradiction checks, freshness checks, and abstention. Track unsupported-claim rate, evidence coverage, abstention rate, stale-source rate, and unauthorized-retrieval blocks against a fixed evaluation set and a fixed system version. Those numbers only help when the denominator and test conditions stay the same.
Add continuous assurance
Next, connect design controls to operating evidence. Map relevant MITRE ATLAS techniques to controls and frozen regression tests. Scan downloaded models before deployment, preserve immutable manifests, separate development and production credentials, and keep releases rollback-capable.
NIST's AI RMF lifecycle of Govern, Map, Measure, and Manage supports continuous controls rather than one-time certification. Use it to shape the control vocabulary and review cycle. Do not treat framework alignment as proof that an implementation is secure.
This sequence works because it reduces blast radius early. Better evaluations matter more after authorization, provenance, and versioning are stable. Monitoring and guardrails still help, but they perform best around an authority-first, fail-closed core.
AletheionAGI provides grounding and evidence-control infrastructure for existing LLMs, RAG pipelines, vector databases, and memory systems, including authorized retrieval, provenance-aware evidence delivery, and abstention when support is insufficient. Visit AletheionAGI to examine how these controls can fit into your AI security architecture before expanding agent authority.



