Agencies generate a high volume of formal communication. Tracking all of it is neither feasible nor useful, so inclusion is deliberately conservative: a statement enters the record as a tracked commitment only if it meets all four criteria below. This produces a smaller, higher-confidence set where every tracked commitment is defensible as a real institutional promise with evaluable specificity.
| Criterion | Requirement |
|---|---|
| Formal public documentation | The commitment appears in an eligible primary source (see Data Sources). Press releases, social media, informal staff statements, journalist-reported claims without agency documentation, and non-public documents are not primary sources. |
| Specificity (3 of 4) | The commitment contains at least three of: an identifiable deliverable, a measurable target, a defined timeframe, and a responsible owner. Meeting all four scores full specificity (Clarity 5); exactly three is tracked with a Missing-Element Flag and a reduced Clarity rating; two or fewer is excluded. |
| Evaluable timeframe | An explicit deadline, plan period, regulatory milestone, or timeframe inferable from document type. Open-ended commitments are not tracked, because delivery cannot be assessed. |
| Agency or office attribution | Attributable to a named agency or office. Diffuse attributions ("the administration," "the federal government") do not qualify alone. Multi-agency commitments are tracked under the lead or signing agency with cross-reference; where neither is identifiable, the commitment is excluded — ownership is not inferred. |
The 3-of-4 specificity threshold is a deliberate trade-off: a stricter 4-of-4 rule would exclude substantial volumes of real commitments that use slightly informal language for one element, while an explicit Missing-Element Flag preserves coverage without hiding the gap.
After a commitment enters the record, its status may be flagged based on later events. Lifecycle flags do not modify delivery; they provide context for interpreting a tracked commitment, and are distinct from the behavioral flags below (which concern how the agency conducted itself).
Each commitment is assessed on four independent axes — Delivery, Difficulty, Impact, and Their Role — carried over unchanged in structure from the Elected Officials methodology. The point maxima match so that agency and official assessments remain comparable. Bucket definitions are adapted to agency action types: rules, guidance, systems, programs, and capabilities.
Delivery measures what the agency actually accomplished relative to what was committed. It is the primary axis and carries the most weight in the final Grade.
| Code | Label | Description | Pts |
|---|---|---|---|
| D4 | Delivered | Commitment fully executed: rule finalized, system operational, program implemented, capability deployed, or stated outcome achieved. | 12 |
| D3 | Substantially Delivered | Most of the commitment executed with documented evidence: rule proposed and on track, system in pilot, program partially operational, or substantial progress toward outcome. | 9 |
| D2 | Partial | Concrete actions taken but commitment not yet executed: rule in early drafting, system in design, program planned but not implemented, or partial outcome achieved. | 6 |
| D1 | Minimal | Planning or announcement only: commitment appears in a strategic plan or regulatory agenda but no observable action toward execution. | 3 |
| D0 | Not Delivered | No meaningful action taken, commitment abandoned, or commitment formally revoked. | 0 |
Negative commitments (a commitment to refrain from an action or to maintain an existing policy) use an inverted scale — D4 (Upheld) through D0 (Violated) — over the full commitment window; the bucket structure is preserved and only the evaluation criteria invert. Phased commitments are assessed against the most recent phase reached, within the stated timeframe.
Their Role modifier (0.0–1.0): the agency's direct causal contribution to the outcome is assessed and applied to Delivery points before aggregation.
Difficulty rewards ambition, but only when delivery occurs. Zero delivery earns zero Difficulty points. The buckets are defined by what kind of action the commitment requires — distinct from the Executive/Legislative/Structural buckets used for elected officials, because agencies face different action structures.
| Code | Label | Description | Max Pts |
|---|---|---|---|
| H3 | Structural | Cross-agency coordination, congressional engagement, court navigation, multi-administration timeline, or significant institutional restructuring. | 5 |
| H2 | Programmatic | Deploy a new system, restructure operations, implement a new program, complete a major capability build, or coordinate with external partners. | 3 |
| H1 | Operational | Publish guidance, issue a technical document, propose a rule under existing authority, complete an internal assessment, or act within an established framework. | 1 |
Impact measures what was at stake. It is not modified by delivery — a failed commitment on a critical issue still carries high impact stakes. Impact is the sum of Scale and Magnitude.
| Scale | Description | Pts | Magnitude | Description | Pts |
|---|---|---|---|---|---|
| S3 | Systemic / National | 4 | M3 | Transformative | 4 |
| S2 | Sectoral / Regional | 2 | M2 | Significant | 2 |
| S1 | Narrow / Targeted | 1 | M1 | Minor / Symbolic | 1 |
Commitment-level Promise Records aggregate into an agency's analytical view. This methodology resists a single agency-level letter Grade: federal agencies are too complex and their missions too varied for one letter to carry meaning, and aggregate grading would obscure the detail the framework is built to surface. Aggregation produces structured agency profiles instead — delivery patterns by commitment type, capability-adjusted patterns, notable cases with evidence trails, and trend analysis across administrations — with Impact weighting the Delivery Ratio at this stage.
"AISI will develop capability sufficient to conduct pre-deployment safety evaluations of frontier AI systems, with an operational capability target of 2025." — NIST AI Safety Institute Strategic Vision (2024)
In isolation, a Promise Record of 6.3 reads as weak performance. The Capability Modifiers below are what let the assessment distinguish that from constrained performance under structural conditions.
Their Role addresses whether the agency had independent authority to deliver, or depended on coordination with other agencies, Congress, the courts, or external actors. A brief rationale is published with each value.
| Value | Situation | Common Examples |
|---|---|---|
| 1.0 | Sole authority. Fully within the agency's regulatory or operational authority; no coordination required. | NIST publication decisions; FDA approvals within existing pathways; internal operational changes. |
| 0.8 | Leads a multi-actor effort. Holds primary authority but directs supporting actors. | Lead-agency rulemakings; agency-directed interagency working groups; agency-directed contractor implementations. |
| 0.6 | Contributes to a broader effort. Significant but not primary authority; depends on coordination with equal or senior actors. | Dual-agency rulemakings; agency contributions to White House initiatives; components within larger federal programs. |
| 0.4 | Dependent on other actors. Defined role but cannot deliver independently; others hold primary or veto authority. | Implementations of congressional appropriations; execution of executive orders where scope is White House–controlled; operations under court oversight. |
| 0.2 | Supporting role. Contributes but has minimal independent authority. | Technical support to another agency's lead effort; participation in industry-led standard-setting; advisory roles. |
Layer 2 answers whether the commitment was delivered. Capability Modifiers add a second question that agency assessment cannot do without: was the agency structurally positioned to deliver? An agency that fails with full resources, clear authority, and political support has demonstrated performance failure; one that fails because Congress rescinded its authorities, the work was deprioritized, contractors collapsed, or the timeline was never achievable has demonstrated structural failure. Conflating the two produces misleading conclusions.
Five factors are each assessed independently on a 1–5 scale (5 fully present · 4 substantially present · 3 mixed · 2 substantially absent · 1 entirely absent), at the commitment level, each with documented evidence. Capability Modifiers do not adjust the Promise Record — that record reflects what was actually delivered. They provide the context for interpreting it. This layer is specific to agencies; the Elected Officials methodology handles equivalent constraints through the Difficulty axis and Their Role instead.
| Factor | Question |
|---|---|
| Resource Adequacy | Did the agency have the budget, staffing, and operational resources to deliver? Assessed against what the specific commitment realistically required, not the agency's overall budget. |
| Authority Alignment | Did the agency have the legal and regulatory authority to deliver independently? Distinct from Their Role, which addresses observed causal contribution; this addresses the prior question of institutional standing to act. |
| Political Support | Did executive-branch leadership and congressional appropriators actively support delivery? |
| Timeline Realism | Was the original deadline achievable given the actions the commitment required? |
| External Dependencies | Did delivery require contractor performance, partner cooperation, court rulings, or other factors outside the agency's control? |
Behavioral flags override or cap standard assessment when the agency's conduct relative to the commitment requires explicit treatment. Each requires a written rationale in the assessment record. Cap-beats-floor precedence applies as in the Elected Officials methodology.
Transparency flags are applied at the assessment level and communicate limitations to readers without altering the result.
Every commitment requires at least two independent sources, and every evidence URL is archived at assessment time to protect against link rot.
Version 2.0 — Initial Departments & Agencies methodology, extending the Elected Officials framework.