Table of Contents

Architecture Decision Records (ADRs)

This folder contains Architecture Decision Records documenting significant technical decisions made in the AgentEval project.

What is an ADR?

An Architecture Decision Record (ADR) captures an important architectural decision along with its context and consequences.

ADR Template

Each ADR follows this structure:

  1. Title - Short descriptive title
  2. Status - Proposed, Accepted, Deprecated, Superseded
  3. Context - The situation and forces that led to this decision
  4. Decision - What we decided to do
  5. Consequences - The results of the decision (positive and negative)
  6. Alternatives Considered - Other options we evaluated

Index

| ADR | Title | Status | Date | |-----|-------|--------|------| | 001 | Metric Naming Prefixes | Proposed | 2026-01-07 | | 002 | Result Directory Structure | Proposed | 2026-01-07 | | 003 | CLI Review Commands | Proposed | 2026-01-07 | | 004 | Trace Recording and Replay | Accepted | 2026-01-07 | | 005 | Model Comparison and stochastic evaluation Architecture | Accepted | 2026-01-08 | | 006 | Service-Based Architecture & DI | Accepted | 2026-01-09 | | 007 | Metrics Taxonomy | Accepted | 2026-01-10 | | 008 | Calibrated Judge for Multi-Model LLM Evaluation | Accepted | 2026-01-12 | | 009 | Benchmark Strategy | Accepted | 2026-01-13 | | 010 | MAF Workflow Integration Architecture | Accepted | 2026-02-14 | | 011 | Workflow Event Processing and Timeout Handling | Accepted | 2026-02-14 | | 012 | Workflow Assertion Design | Accepted | 2026-02-14 | | 013 | Microsoft Agent Framework RC1 Upgrade | Accepted | 2026-02 | | 014 | Dataset Pipeline — Two-Model Architecture | Accepted | 2026-02-24 | | 015 | Extension Registration — Manual vs Auto-Discovery | Accepted | 2026-02-25 | | 016 | Monolith Modularization | Accepted | 2026-02-26 | | 017 | Unified Benchmarks Namespace (Convention 1-4) | Implemented (v0.10.0-beta) — Convention 2 is superseded by ADR-032, accepted 2026-09-07; Convention 3 stands as a catalog | 2026-05-17 | | 018 | Compliance.Core and Cross-Cutting Shared Extractions | Accepted | 2026-05-31 | | 019 | Chat-Boundary Tracing and the Two-Layer Recording Model (Glass Box) | Accepted | 2026-05-31 | | 020 | AgentTrace v1.1 Schema (Glass Box additive fields) | Accepted | 2026-05-31 | | 021 | Judge-Primary Grading for Semantic Red-Team Oracles (RedTeam Phase B) | Accepted (extended by 022) | 2026-06-18 | | 022 | Grading by Decomposition (Composite Sub-Evaluators) — RedTeam Phase C | Accepted (extends 021; B.3 executed) | 2026-06-21 | | 023 | Decompose the Misinformation Oracle (Confabulation ⊕ Existence-Denial) | Accepted | 2026-06-22 | | 024 | Split-then-Gate Decomposition (Gated Trees) and Its Bounds | Accepted | 2026-06-23 | | 025 | Gatekeeper — Runtime Fail-Closed Enforcement Middleware | Accepted | 2026-07-05 | | 026 | TypedMemEval — A Mechanism-Isolating Memory Benchmark Family | Accepted (implemented v0.22.0-beta) | 2026-08-15 | | 027 | TypedMemEval — Semantic, Temporal and Bitemporal Verticals (design) | Proposed (design only; generation gated) | 2026-08-18 | | 028 | TypedMemEval — accept a shape on measured discrimination, not a coverage proxy | Proposed | 2026-08-31 | | 029 | Procedural memory — a finding, and a request to reopen a bilateral agreement | Accepted, then REVERSED by its own §10–§12: the Procedural vertical is built (80q, headroom +0.81) and released standalone | 2026-09-03 | | 030 | Meta-evaluation is the lane; contract unification is not — chance floors, exact tests, applicability as first-class; IEval kept by adaptation; deterministic assertion catalogue rejected by name | Accepted (Slice 0 shipped in a396c5b4; Slices 1–2 shipped; §11 amendment 2026-09-07 records the AE-04 join, the 79 / 7 / 1 census and 15 corrections of record; Q4(ii), Q5, Q6 remain the owner's) | 2026-09-05 | | 031 | Eval Packs — adversarial verdict DON'T BUILD as scoped; SHIP REDUCED to five items on the existing subjects/ tree, no new format | Rejected as scoped; SHIP REDUCED — S2 and S5 shipped, S3 half-shipped (writer half = Q4(ii)), S1 deferred, S4 CLOSED by Q5’s answer (2026-09-07): defer the API, a degraded BenchmarkArm is the control (§0.1; §12 amendment 2026-09-07) | 2026-09-05 | | 032 | Benchmarks are definitions; runs bind subjects; scores are meta — aggregation stops demanding an IEval per weight (four stubs deleted); the floor door refuses composites; a deterministic benchmark contract over AE-04's join, IOutputStore and agenteval compare unchanged; no new verb, package or schema field | Accepted (2026-09-07) — Q-A answered INSIDE the rule and Wave 2 BUILT (definition records, BenchmarkScore, BenchmarkRunner, the EvalJoin/02 sample, one in-repo benchmark converted); Q4(ii)/Q5 deferred and Q6 answered yes-but-staged, each refused in code | 2026-09-07 |

Template based on Michael Nygard's ADR format