Sources and retained corrections
This is the supplied research register, preserved separately from the newer source-review layer. A citation is not runtime verification. Review observations retain unresolved access and support limits.
Current review differences
These findings qualify the retained register below. Original input bytes and historical inspection statements remain unchanged.
- AGENTEVAL26 (Table 8, PDF page 11): Caption and abstract say 21 subcategories; visible Level 3 rows total 20. No missing row inferred.
- RIFL15 (Current body at supplied PDF URL): The retrieved artifact is a research paper, not a slide presentation. Its first page includes abstract, authors and SOSP15 copyright notice.
Remaining review: Full claim-by-claim support for all 283 mappings is not independently established. Abstract-only, selected-page, landing-page and truncated-source limitations remain explicit; no source is a runtime prevention certificate.
EFFECT26
Trofimov & Novikov, When Tool Calls Succeed but Workflows Fail (2609.15397v1)
External-effect history, A1-A8 and interface-dependent guarantee boundaries. Not a proof that a particular deployed harness implements the prerequisites.
Type: Research preprint. Supplied access date: 2026-09-22.
RAJ26
Raj et al., Model or Harness? (2607.28802v1)
Interaction-centric, role-specific fault classification. Illustrative reports are not a population prevalence estimate. This is not the Taraghi paper.
Type: Research preprint. Supplied access date: 2026-09-22.
MCPFAULT26
Taraghi, Morovati & Khomh, Real Faults in MCP Software (2603.05637v1)
MCP-specific software defects, including setup, discovery, connections, tools, host integration and diagnostics. Counts describe the sampled issue corpus.
Type: Empirical research preprint. Supplied access date: 2026-09-22.
AGFAULT26
Shah et al., Characterizing Faults in Agentic AI (2603.06847v1)
Agent software fault types, symptoms and causes are separate taxonomies. PDF contains template publication metadata; cited as an arXiv preprint, not as the journal implied by that footer.
Type: Empirical research preprint. Supplied access date: 2026-09-22.
MAST25
Cemri et al., Why Do Multi-Agent LLM Systems Fail? (2503.13657v3)
Fourteen coordination failure labels. Multi-agent evidence informs boundary analysis; it does not imply Camden should introduce multiple reasoning executors.
Type: Research paper. Supplied access date: 2026-09-22.
TOOLSCAN25
Kokane et al., ToolScan (2411.13547v2)
Tool-call error classes, including missing/repeated calls, argument name/type/value, function identity and formatting. These are not all external-effect failures.
Type: Research paper. Supplied access date: 2026-09-22.
TOOLBENCHX26
Beyond Function Calling: ToolBench-X (2606.25819v1)
Five tool-environment hazard families in controlled, recoverable fixtures; do not generalize fixture recoverability to arbitrary providers.
Type: Research preprint. Supplied access date: 2026-09-22.
CHAOS26
Zhang et al., When Agentic Executions Fail / AgentChaosBench (2608.14680v1)
Ten injected runtime faults, plus no-fault controls. Injection families and trace-localization scores are not a complete incident ontology.
Type: Benchmark preprint. Supplied access date: 2026-09-22.
MEMFAIL26
Garg et al., MemFail (2605.26667v1)
Separates summary, storage, retrieval and downstream reasoning failures. Includes coexistence and conditional facts; synthetic benchmark limits apply.
Type: Benchmark preprint. Supplied access date: 2026-09-22.
MEMDRIFT26
Memory-Induced Tool-Drift in LLM Agents (2605.24941v1)
Stored personal tendencies can affect tool parameters outside relevant scope. Effect sizes are benchmark-specific.
Type: Research preprint. Supplied access date: 2026-09-22.
MINJA25
Dong et al., Memory Injection Attacks via Query-Only Interaction (2503.03704v4)
Query-only memory poisoning mechanism. Specific attack success rates do not apply to every memory architecture.
Type: Security research. Supplied access date: 2026-09-22.
MEMRISK26
Al-Tawaha et al., Remembering More, Risking More (2605.17830v1)
Longitudinal safety risks in memory-equipped agents. Not evidence for every compression-laundering story in the supplied draft.
Type: Research preprint; abstract checked. Supplied access date: 2026-09-22.
PLAUSIBLE26
When Errors Become Narratives (2606.14589v1)
Silent failures and fluent narratives that obscure them. Case-study findings are not universal failure frequencies.
Type: Production-runtime case study preprint. Supplied access date: 2026-09-22.
TOCTOU25
Lilienthal & Hong, Mind the Gap (2508.17155)
Check/use races in agent tool workflows. Atomic or conditional operations are still subject to the actual target contract.
Type: Security research; abstract checked. Supplied access date: 2026-09-22.
PROGRESS26
Liu et al., Ask the Tool, Do Not Guess (2609.18849v1)
Tool-progress signals for serving/cache decisions. Does not establish cross-tenant KV corruption or semantic state splicing.
Type: Systems preprint; abstract checked. Supplied access date: 2026-09-22.
AGENTLEAK26
El Yagoubi et al., AgentLeak (2602.11510v1)
Seven disclosure channels and six attack families. Coverage and evaluation strength differ across channels; a trust-boundary violation is not automatically a public breach.
Type: Privacy benchmark preprint. Supplied access date: 2026-09-22.
AGENTRX26
Barke et al., AgentRx (2602.02475v1)
Failure localization through constraints and execution evidence. An inferred critical step is an analysis result, not omniscient causal ground truth.
Type: Research preprint. Supplied access date: 2026-09-22.
AGENTEVAL26
Guo et al., AgentEval (2604.23581v1)
Step/dependency-aware evaluation and hierarchical categories. DAG assumptions and judge calibration constrain applicability. Table 8 in v1 visibly enumerates 20 Level-3 rows although its caption states 21; the table was checked in the PDF image. The crosswalk uses the 20 visible rows and does not invent a twenty-first.
Type: Research preprint. Supplied access date: 2026-09-22.
Current review qualification: Caption and abstract say 21 subcategories; visible Level 3 rows total 20. No missing row inferred.
AIRT26
Microsoft AI Red Team, Taxonomy of Failure Modes in Agentic AI Systems v2.0
Thirty-four safety/security labels; prospective threats are separately marked by the document. It is not a certification checklist.
Type: First-party security taxonomy, April 2026 PDF. Supplied access date: 2026-09-22.
OWASP26
OWASP Top 10 for Agentic Applications for 2026
Ten umbrella risk classes. Umbrella labels are crosswalks, not ten additional non-overlapping mechanisms. Official launch article independently enumerates ASI01-ASI10. The resource-page download click failed; no claim of full PDF retrieval is made.
Type: Security taxonomy. Supplied access date: 2026-09-22.
HARNESSES25
Anthropic, Effective harnesses for long-running agents
Session continuation, progress artifacts and end-to-end testing. Not proof of external provider delivery or universal model independence.
Type: First-party engineering research, 2025-11-26. Supplied access date: 2026-09-22.
AWSID
AWS Builders Library, Making retries safe with idempotent APIs
Intent identity, parameter binding, atomic deduplication, late requests. Identical parameters need not mean identical intent.
Type: First-party engineering. Supplied access date: 2026-09-22.
STRIPEID
Stripe API, Idempotent requests
Key retention, parameter comparison and cached outcomes for this documented API surface. Do not conflate v1 and v2 semantics.
Type: Provider contract. Supplied access date: 2026-09-22.
STRIPEV2
Different API version, method and scope constraints; v2 documentation describes a different deduplication interval from v1.
Type: Provider contract. Supplied access date: 2026-09-22.
OUTBOX
AWS Prescriptive Guidance, Transactional outbox
Atomic local data/outbox commit does not eliminate duplicate relay delivery or supply remote transaction participation.
Type: First-party architecture guidance. Supplied access date: 2026-09-22.
SAGA
Microsoft Azure, Compensating Transaction pattern
Business compensation need not restore exact initial state or follow reverse order. Compensation itself can fail and must account for concurrent changes.
Type: First-party architecture guidance. Supplied access date: 2026-09-22.
PGISO
PostgreSQL 18, Transaction Isolation
Isolation anomalies and serializable transaction retry requirements. Database transaction guarantees do not automatically span external tools.
Type: Database documentation. Supplied access date: 2026-09-22.
PGPITR
PostgreSQL 18, Continuous Archiving and PITR
Recovery history, archive continuity and timelines. Restoring a database does not restore external effects to the same point in time.
Type: Database documentation. Supplied access date: 2026-09-22.
SQLITE
SQLite, How To Corrupt An SQLite Database File
Durability, journals, locking, backup pairing, storage and memory faults. Historical implementation incidents are not assertions that every current version has them.
Type: First-party failure documentation. Supplied access date: 2026-09-22.
TEMPORAL
Replay determinism and version compatibility. Workflow replay and external activity effects are different boundaries.
Type: Durable workflow documentation. Supplied access date: 2026-09-22.
FENCES
Kleppmann, How to do distributed locking
Lease expiry, process pauses and target-enforced fencing. A client-side lock cannot fence a target that ignores its token.
Type: Primary technical analysis. Supplied access date: 2026-09-22.
FLP85
Fischer, Lynch & Paterson, Impossibility of Distributed Consensus with One Faulty Process
Termination limit under its asynchronous deterministic consensus model; not a universal statement that practical systems cannot recover.
Type: Formal research paper, 1985. Supplied access date: 2026-09-22.
RIFL15
Lee et al., Implementing Linearizability at Large Scale and Low Latency
Completion records, durable duplicate handling and record lifetime. Cited as the conference presentation, not mistaken for the full proceedings paper.
Type: Authors conference presentation, SOSP 2015. Supplied access date: 2026-09-22.
Current review qualification: The retrieved artifact is a research paper, not a slide presentation. Its first page includes abstract, authors and SOSP15 copyright notice.
RFC9110
Method semantics, conditionals, status codes and retries. HTTP success has the meaning supplied by the endpoint, not a universal business-completion meaning.
Type: Standard, 2022. Supplied access date: 2026-09-22.
RFC9700
IETF RFC 9700, OAuth 2.0 Security Best Current Practice
Token audience, redirect validation, authorization binding and refresh security. Possession of credentials does not supply business authorization.
Type: Security standard, 2025. Supplied access date: 2026-09-22.
RFC8785
IETF RFC 8785, JSON Canonicalization Scheme
Canonical bytes for signing/hashing. Does not establish truth, authorization, completeness or harmlessness of those bytes.
Type: Canonicalization specification, 2020. Supplied access date: 2026-09-22.
RFC3339
IETF RFC 3339, Date and Time on the Internet
Instant representation and offsets; future civil recurrences still need zone and recurrence semantics.
Type: Time representation standard. Supplied access date: 2026-09-22.
K8SAPI
Versioned reads, conditional writes, watches and pagination. Contracts must be checked for the actual server version.
Type: Control-plane documentation. Supplied access date: 2026-09-22.
K8SGC
Kubernetes, Garbage Collection
Owners, dependent resources and cleanup boundaries. Useful precedent for orphan and incorrectly collected work artifacts.
Type: Control-plane documentation. Supplied access date: 2026-09-22.
K8SCTRL
Desired/observed state reconciliation; declarative state alone does not prove a live controller or converged outcome.
Type: Control-plane documentation. Supplied access date: 2026-09-22.
SRELOAD
Google SRE Book, Handling Overload
Admission, overload, queueing and load shedding. Model resources and recovery work participate in the same capacity constraints.
Type: First-party systems engineering. Supplied access date: 2026-09-22.
SRECASCADE
Google SRE Book, Addressing Cascading Failures
Retry amplification, shared dependency failures and recovery load. Independent-looking components can share a fault domain.
Type: First-party systems engineering. Supplied access date: 2026-09-22.
MCPSEC
MCP Security Best Practices, 2025-11-25 documentation
Confused deputy, token passthrough, SSRF, sessions and scope. The protocol does not automatically enforce every host security requirement.
Type: Protocol security guidance. Supplied access date: 2026-09-22.
A2A
A2A Protocol specification, v1.0-era page retrieved 2026-09-22
Task states, identities, streams, capabilities and discovery. The retrieved page uses /.well-known/agent-card.json; signatures are optional and are not execution grants.
Type: Protocol specification snapshot. Supplied access date: 2026-09-22.
GHSEC
GitHub Actions, Secure use reference
Untrusted contributions, workflow privilege and self-hosted runner risks. Public documentation discovery is separate from executing public code.
Type: First-party security guidance. Supplied access date: 2026-09-22.
SLSA
SLSA v1.1, Supply chain threats
Threats across source, build, dependencies and artifact distribution. Provenance narrows claims; it does not certify semantic correctness.
Type: Supply-chain security specification. Supplied access date: 2026-09-22.
CWE78
MITRE CWE-78, OS Command Injection
Untrusted data crossing into shell syntax. Defensive mechanism classification only; catalog examples contain no operational attack payload.
Type: Weakness definition. Supplied access date: 2026-09-22.
CWE918
MITRE CWE-918, Server-Side Request Forgery
Untrusted destinations reaching unintended services. Applies to tools, callbacks, document fetches and registries.
Type: Weakness definition. Supplied access date: 2026-09-22.
CWE400
MITRE CWE-400, Uncontrolled Resource Consumption
Compute, memory, storage and network exhaustion. Each concrete resource boundary still needs its own measurement and enforcement.
Type: Weakness definition. Supplied access date: 2026-09-22.
TLA
Specification and model-checking foundation. The existence of a model or proof does not establish implementation refinement or environment assumptions.
Type: Author-hosted formal methods book. Supplied access date: 2026-09-22.
DELTABOX
Dong et al., DeltaBox (2605.22781)
Checkpoint/rollback of sandbox process and file state. Does not promise rollback of externally observed or physical effects.
Type: Systems preprint; abstract checked. Supplied access date: 2026-09-22.
CONSENSUS26
Rodrigues, Hallucination as Context Drift (2606.21666v1)
A broadcast-contamination result in travel scenarios did not replicate in the software domain. Not a universal numerical penalty for synchronization.
Type: Small controlled-study preprint; abstract checked. Supplied access date: 2026-09-22.
AGENTDOJO24
Debenedetti et al., AgentDojo (2406.13352)
Prompt-injection evaluation with tool use and utility tradeoffs. No defense in this catalog is claimed universally injection-proof.
Type: Security benchmark. Supplied access date: 2026-09-22.
VERIFIED26
Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures (2608.02645)
Timeout, visibility, partial-operation and stale-conflict fixtures. Verification policy remains conditional on target capabilities.
Type: Research preprint; abstract checked. Supplied access date: 2026-09-22.
SYNTHESIS26
Albayaydh et al., Beyond the Leaderboard (2607.05775)
Cross-benchmark grouping and measurement-validity concerns. Full text was not retrieved; no fine-grained incident claims are taken from it.
Type: Literature-synthesis preprint; abstract checked. Supplied access date: 2026-09-22.
RFC6973
RFC 6973, Privacy Considerations for Internet Protocols
Data minimization, disclosure, correlation, secondary use and retention. Not jurisdiction-specific legal advice or a certification of compliance.
Type: Informational RFC. Supplied access date: 2026-09-22.
NASAVV
NASA Systems Engineering Handbook, Verification and Validation Plan Outline
Separates requirements verification, intended-use validation and integrated-system evidence. Used as a precedent for original boundary examples, not as an agent incident report.
Type: First-party systems engineering guidance. Supplied access date: 2026-09-22.
The 22 supplied corrections
These original corrections retain the source packet's stated inspection scope. Current review observations remain distinct.
1. 41-mode paper authorship
The supplied draft attributes Model or Harness? to Taraghi. The inspected paper is by Raj and colleagues. Taraghi and colleagues wrote the separate MCP-software-fault study.
Sources: RAJ26, MCPFAULT26.
2. A6 definition
Dependency on a nonsurviving effect is distinct from merely discarding a speculative branch. Both are represented, without pretending the two labels are identical.
Sources: EFFECT26.
3. Outbox scope
An outbox can atomically persist a local business change and local event intent. Its relay can still redeliver, and it does not add outcome visibility or atomic participation to an opaque external provider.
Sources: OUTBOX, EFFECT26.
4. Idempotency identity
A key may be randomly generated once and durably retained. Deterministic content-only derivation is not required and can collapse two legitimate identical requests. Scope, payload binding, retention and atomic target handling matter.
Sources: AWSID, STRIPEID, STRIPEV2.
5. Absence is not quiescence
A current absence read, even if fresh, does not prove a previously queued request cannot execute later. Safe retry needs the relevant target-side guarantee or positively established termination of the earlier attempt.
Sources: AWSID, FENCES.
6. Saga versus atomic commit
A saga is a deliberate compensation model, not simply a broken two-phase commit. Compensation may restore a valid business state without restoring the exact old state, and must account for intervening legitimate changes.
Sources: SAGA, EFFECT26.
7. Compensating under caller uncertainty
Unconditional undo under an unknown outcome is unsafe. A target-supported conditional compensation primitive can legitimately resolve that condition atomically even if the caller does not first know the outcome.
Sources: SAGA, EFFECT26.
8. Fencing location
A gateway can reject future stale admissions; it cannot retract an already forwarded, uncontrollable remote command. The enforcement point and in-flight set must be named.
Sources: FENCES, EFFECT26.
9. Heartbeat interpretation
Missing heartbeat means suspected failure, not proof of death. A live heartbeat also does not prove mission progress.
Sources: FLP85, SRECASCADE.
10. Verifier semantics
Neither a second LLM nor a deterministic checker is automatically an independent truth oracle. Verification needs appropriate evidence, a valid predicate and protection from common-mode or subject-controlled failure.
Sources: AGENTRX26, AGENTEVAL26, NASAVV.
11. Receipts and cryptography
A signature or hash can establish origin/integrity within a trust model, not the truth, completeness, freshness or physical meaning of the signed assertion. Provider signatures are not universally required when other trustworthy observation paths suffice.
Sources: SLSA, RFC9700.
12. Isolation and prompt injection
A sandbox label, data tag or signed tool description is not a universal injection or containment guarantee. Actual capabilities, mounts, egress and trusted enforcement still matter.
Sources: MCPSEC, GHSEC, AGENTDOJO24.
13. HTTP error classes
Error handling must use the actual provider contract. A timeout or some server error responses may follow an applied effect; some authentication failures are legitimately refreshable. Blanket retry/fatal classifications are not sound.
Sources: RFC9110, RFC9700, STRIPEID.
14. Audit retention
Durable lineage does not require retaining every private reasoning token forever. It requires sufficient scoped intent, decisions, constraints, effects, evidence, corrections and continuation records under an explicit retention policy.
Sources: RFC6973, PGPITR.
15. Coverage denominator
Not every task can enumerate all future work up front. Snapshot, watermark, pagination or another justified dynamic closure rule may be needed. A success count alone does not prove correct set membership.
Sources: K8SAPI, AGENTEVAL26.
16. KV-cache extrapolation
Ask the Tool, Do Not Guess supports tool-progress and serving-resource analysis. It does not establish the cross-tenant corruption or prefix-splicing incidents asserted in the supplied draft.
Sources: PROGRESS26.
17. Broadcast percentage extrapolation
The cited context-drift study explicitly says its software-domain follow-up did not replicate the travel-domain broadcast degradation. The number is not a universal property of agent communication.
Sources: CONSENSUS26.
18. Visual input distinction
DOM-hidden text, tiny rendered text and pixels absent from a raster are different channels. A screenshot model cannot read text that is genuinely absent from its pixel input.
Sources: AIRT26.
19. A2A discovery path
The retrieved specification page uses /.well-known/agent-card.json. The older agent.json path in the supplied draft should not be presented as the checked current path. Discovery metadata is not an authority grant.
Sources: A2A.
20. AgentEval count discrepancy
The v1 Table 8 caption states 21 Level-3 categories, but the inspected PDF table shows 20 named rows. This catalog maps the visible rows and does not manufacture a twenty-first.
Sources: AGENTEVAL26.
21. Publication metadata
Characterizing Faults in Agentic AI contains an inconsistent 2018 journal-template footer. It is cited here as the retrieved arXiv preprint, not as a verified 2018 journal publication.
Sources: AGFAULT26.
22. RIFL representation
The retrieved SOSP item is the authors' slide deck. It is not silently represented as a full-text proceedings-paper review.
Sources: RIFL15.