Nonpartisan Government Accountability
PolicyLogic
How We Apply Our Methodology
INDEPENDENT & TRANSPARENT METHODOLOGY
PolicyLogic
Home
Methodology
Elected Officials Depts & Agencies Presidential & International AI Pipeline Limitations Contact
Elected Officials Departments & Agencies Presidential & International AI Pipeline Limitations
Departments & Agencies Methodology
How PolicyLogic assesses federal departments and agencies on the commitments they make in formal public documents. Version 2.0 — operative standard. Extends the Elected Officials methodology; all prior versions superseded.
How assessments are produced. Assessments are produced by an automated pipeline and verified by researchers against the public record using documented procedures. Verification is ongoing rather than complete. Each assessment is labeled with its current state.
Relationship to the Elected Officials methodology. This methodology extends the Elected Officials framework to federal departments and agencies. The two share the same four-axis structure (Delivery, Difficulty, Impact, Their Role), the same evidence tiers, and the same pipeline. They diverge where agency commitments require different inclusion rules, source types, and one added context layer. The shared question anchors both: did they deliver, was it hard, did it matter?

What Counts as a Commitment

Agencies generate a high volume of formal communication. Tracking all of it is neither feasible nor useful, so inclusion is deliberately conservative: a statement enters the record as a tracked commitment only if it meets all four criteria below. This produces a smaller, higher-confidence set where every tracked commitment is defensible as a real institutional promise with evaluable specificity.

CriterionRequirement
Formal public documentationThe commitment appears in an eligible primary source (see Data Sources). Press releases, social media, informal staff statements, journalist-reported claims without agency documentation, and non-public documents are not primary sources.
Specificity (3 of 4)The commitment contains at least three of: an identifiable deliverable, a measurable target, a defined timeframe, and a responsible owner. Meeting all four scores full specificity (Clarity 5); exactly three is tracked with a Missing-Element Flag and a reduced Clarity rating; two or fewer is excluded.
Evaluable timeframeAn explicit deadline, plan period, regulatory milestone, or timeframe inferable from document type. Open-ended commitments are not tracked, because delivery cannot be assessed.
Agency or office attributionAttributable to a named agency or office. Diffuse attributions ("the administration," "the federal government") do not qualify alone. Multi-agency commitments are tracked under the lead or signing agency with cross-reference; where neither is identifiable, the commitment is excluded — ownership is not inferred.

The 3-of-4 specificity threshold is a deliberate trade-off: a stricter 4-of-4 rule would exclude substantial volumes of real commitments that use slightly informal language for one element, while an explicit Missing-Element Flag preserves coverage without hiding the gap.

Commitment Lifecycle Flags

After a commitment enters the record, its status may be flagged based on later events. Lifecycle flags do not modify delivery; they provide context for interpreting a tracked commitment, and are distinct from the behavioral flags below (which concern how the agency conducted itself).

Revoked
Explicitly rescinded by the agency or a superseding authority. Original record preserved; delivery assessment stops at the revocation date.
Superseded
Replaced by a later commitment from the same agency that materially alters the deliverable, target, timeframe, or scope. Original preserved with a reference to the replacement.
Contingent
Execution depends on action by another actor not yet taken — most commonly congressional appropriation. Not assessed for delivery until the contingency resolves.
Awaiting Translation
An executive order or OMB memorandum directed at the agency that it has not yet translated into a formal commitment. Surfaced separately, not scored as failure; elapsed time is shown.
Lapsed
A cross-administration commitment dropped without explicit revocation, applied when a new administration does not formally restate it within 12 months of transition.
Historical
Entered retroactively from a period preceding the record's start date; flagged and dated.

Assessment Framework

Each commitment is assessed on four independent axes — Delivery, Difficulty, Impact, and Their Role — carried over unchanged in structure from the Elected Officials methodology. The point maxima match so that agency and official assessments remain comparable. Bucket definitions are adapted to agency action types: rules, guidance, systems, programs, and capabilities.

Axis 1 — Delivery (max 12 points)

Delivery measures what the agency actually accomplished relative to what was committed. It is the primary axis and carries the most weight in the final Grade.

CodeLabelDescriptionPts
D4DeliveredCommitment fully executed: rule finalized, system operational, program implemented, capability deployed, or stated outcome achieved.12
D3Substantially DeliveredMost of the commitment executed with documented evidence: rule proposed and on track, system in pilot, program partially operational, or substantial progress toward outcome.9
D2PartialConcrete actions taken but commitment not yet executed: rule in early drafting, system in design, program planned but not implemented, or partial outcome achieved.6
D1MinimalPlanning or announcement only: commitment appears in a strategic plan or regulatory agenda but no observable action toward execution.3
D0Not DeliveredNo meaningful action taken, commitment abandoned, or commitment formally revoked.0

Negative commitments (a commitment to refrain from an action or to maintain an existing policy) use an inverted scale — D4 (Upheld) through D0 (Violated) — over the full commitment window; the bucket structure is preserved and only the evaluation criteria invert. Phased commitments are assessed against the most recent phase reached, within the stated timeframe.

Their Role modifier (0.0–1.0): the agency's direct causal contribution to the outcome is assessed and applied to Delivery points before aggregation.

Adjusted Delivery = Delivery Points × Their Role

Axis 2 — Difficulty (max 5 points, earned proportionally)

Difficulty rewards ambition, but only when delivery occurs. Zero delivery earns zero Difficulty points. The buckets are defined by what kind of action the commitment requires — distinct from the Executive/Legislative/Structural buckets used for elected officials, because agencies face different action structures.

CodeLabelDescriptionMax Pts
H3StructuralCross-agency coordination, congressional engagement, court navigation, multi-administration timeline, or significant institutional restructuring.5
H2ProgrammaticDeploy a new system, restructure operations, implement a new program, complete a major capability build, or coordinate with external partners.3
H1OperationalPublish guidance, issue a technical document, propose a rule under existing authority, complete an internal assessment, or act within an established framework.1
Difficulty Earned = Bucket Points × (Delivery Points / 12)
Difficulty Earned scales with unmodified Delivery Points, not Adjusted Delivery. Their Role discounts the Delivery axis only; it does not compound into the Difficulty axis.
Example: H3 commitment with D2 delivery = 5 × (6/12) = 2.5 pts · H3 with D0 = 0 pts

Axis 3 — Impact (max 8 points)

Impact measures what was at stake. It is not modified by delivery — a failed commitment on a critical issue still carries high impact stakes. Impact is the sum of Scale and Magnitude.

ScaleDescriptionPtsMagnitudeDescriptionPts
S3Systemic / National4M3Transformative4
S2Sectoral / Regional2M2Significant2
S1Narrow / Targeted1M1Minor / Symbolic1
Specificity Cap (Clarity Rule): a commitment at Clarity 2 or lower has Magnitude capped at M1. An agency cannot receive high Impact credit for a commitment that was never specific enough to produce a measurable outcome.

Promise Record

Promise Record = Adjusted Delivery + Difficulty Earned
Maximum per commitment: 12 (Adjusted Delivery) + 5 (Difficulty Earned) = 17 points. Impact is excluded from the Promise Record because it re-enters as a weight on the Delivery Ratio during aggregation; including it in both would double-count it.
Note on comparability with the Elected Officials methodology. There, the Promise Record is out of 25 and includes Impact directly, because that methodology does not perform the agency-level Impact-weighted aggregation described here. The two Records answer the same question but use different denominators by design. Delivery and Impact points are identical across both; only where Impact is applied differs.

Assessment Calculation

Commitment-level Promise Records aggregate into an agency's analytical view. This methodology resists a single agency-level letter Grade: federal agencies are too complex and their missions too varied for one letter to carry meaning, and aggregate grading would obscure the detail the framework is built to surface. Aggregation produces structured agency profiles instead — delivery patterns by commitment type, capability-adjusted patterns, notable cases with evidence trails, and trend analysis across administrations — with Impact weighting the Delivery Ratio at this stage.

Worked Example

Worked Example · AI Safety Capability · NIST (Programmatic, H2)

"AISI will develop capability sufficient to conduct pre-deployment safety evaluations of frontier AI systems, with an operational capability target of 2025." — NIST AI Safety Institute Strategic Vision (2024)

InclusionTracked — 3 of 4 specificity elements (deliverable, timeframe, owner met; measurable target absent). Missing-Element Flag applied; Clarity 4.
Delivery BucketD2 — institute established, initial partnerships in place, but pre-deployment evaluation not operational at the implied scope
Delivery Points6
Their Role0.8 — leads a multi-actor effort; coordinates with industry, other agencies, and international counterparts
Adjusted Delivery6 × 0.8 = 4.8 pts
DifficultyH2 (Programmatic) = 3 max pts
Difficulty Earned3 × (6/12) = 1.5 pts
Impact (separate)S3 + M3 = 4 + 4 = 8 pts — applied in aggregation, not in the Promise Record
Promise Record4.8 + 1.5 = 6.3 / 17

In isolation, a Promise Record of 6.3 reads as weak performance. The Capability Modifiers below are what let the assessment distinguish that from constrained performance under structural conditions.

Their Role — Lookup Table

Their Role addresses whether the agency had independent authority to deliver, or depended on coordination with other agencies, Congress, the courts, or external actors. A brief rationale is published with each value.

ValueSituationCommon Examples
1.0Sole authority. Fully within the agency's regulatory or operational authority; no coordination required.NIST publication decisions; FDA approvals within existing pathways; internal operational changes.
0.8Leads a multi-actor effort. Holds primary authority but directs supporting actors.Lead-agency rulemakings; agency-directed interagency working groups; agency-directed contractor implementations.
0.6Contributes to a broader effort. Significant but not primary authority; depends on coordination with equal or senior actors.Dual-agency rulemakings; agency contributions to White House initiatives; components within larger federal programs.
0.4Dependent on other actors. Defined role but cannot deliver independently; others hold primary or veto authority.Implementations of congressional appropriations; execution of executive orders where scope is White House–controlled; operations under court oversight.
0.2Supporting role. Contributes but has minimal independent authority.Technical support to another agency's lead effort; participation in industry-led standard-setting; advisory roles.

Capability Modifiers

Layer 2 answers whether the commitment was delivered. Capability Modifiers add a second question that agency assessment cannot do without: was the agency structurally positioned to deliver? An agency that fails with full resources, clear authority, and political support has demonstrated performance failure; one that fails because Congress rescinded its authorities, the work was deprioritized, contractors collapsed, or the timeline was never achievable has demonstrated structural failure. Conflating the two produces misleading conclusions.

Five factors are each assessed independently on a 1–5 scale (5 fully present · 4 substantially present · 3 mixed · 2 substantially absent · 1 entirely absent), at the commitment level, each with documented evidence. Capability Modifiers do not adjust the Promise Record — that record reflects what was actually delivered. They provide the context for interpreting it. This layer is specific to agencies; the Elected Officials methodology handles equivalent constraints through the Difficulty axis and Their Role instead.

FactorQuestion
Resource AdequacyDid the agency have the budget, staffing, and operational resources to deliver? Assessed against what the specific commitment realistically required, not the agency's overall budget.
Authority AlignmentDid the agency have the legal and regulatory authority to deliver independently? Distinct from Their Role, which addresses observed causal contribution; this addresses the prior question of institutional standing to act.
Political SupportDid executive-branch leadership and congressional appropriators actively support delivery?
Timeline RealismWas the original deadline achievable given the actions the commitment required?
External DependenciesDid delivery require contractor performance, partner cooperation, court rulings, or other factors outside the agency's control?
Pattern reading. Low delivery combined with low Timeline Realism and Resource Adequacy but high Authority Alignment reads as constrained performance under structural conditions — not weak performance. Surfacing that distinction is the purpose of the layer, and it drives the Outlier-review tier, where a distinctive capability pattern triggers human verification before publication.

Behavioral Flags

Behavioral flags override or cap standard assessment when the agency's conduct relative to the commitment requires explicit treatment. Each requires a written rationale in the assessment record. Cap-beats-floor precedence applies as in the Elected Officials methodology.

Reversed
Delivery = D0
Agency actively undid a commitment it previously delivered. The reversal enters the record as a new commitment.
Redefined
Capped at D2
Definition of success changed mid-commitment through informal means rather than a new formal source. Reflects delivery against shifted goalposts.
Externally Blocked
D2 minimum if pursued
A documented external obstacle prevented delivery despite agency effort — rescinded authorities, injunctions, OMB intervention, contractor failure outside agency control.
Credit Overclaimed
Their Role capped at 0.4
Agency claimed credit for outcomes driven mainly by predecessors, other agencies, or external conditions.
Contradictory
Links the pair
Commitment conflicts with another tracked commitment from the same agency. Both are assessed; the flag links them.
Self-Reported Only
Capped at D2
Delivery evidence comes entirely from agency self-reporting without independent verification. Agency-specific, because agency self-reporting is far more common than for elected officials.

Transparency Flags

Transparency flags are applied at the assessment level and communicate limitations to readers without altering the result.

Contested Evidence
Sources disagree on whether the commitment was delivered. Assessed conservatively (the lower bucket) with the conflict noted.
Insufficient Evidence
Fewer than two independent sources. Displayed but excluded from agency-level aggregation; the evidence gap is indicated.
Provisional
Produced by the automated pipeline; researcher verification against the public record is ongoing.
Under Review
Classification actively being re-evaluated due to new evidence.
Independent Agency
Commitment from an independent regulatory body (Federal Reserve, FCC, SEC, FTC, NLRB, and others), handled separately because its accountability structure differs from executive-branch agencies.
Confidence Band
High / Medium / Low rating reflecting evidence strength and application certainty for the specific commitment.

Data Sources

Every commitment requires at least two independent sources, and every evidence URL is archived at assessment time to protect against link rot.

Changelog

Version 2.0 — Initial Departments & Agencies methodology, extending the Elected Officials framework.