Nonpartisan Government Accountability
PolicyLogic
How We Apply Our Methodology
INDEPENDENT & TRANSPARENT METHODOLOGY
PolicyLogic
Home
Methodology
Elected Officials Depts & Agencies Presidential & International AI Pipeline Limitations Contact
Elected Officials Departments & Agencies Presidential & International AI Pipeline Limitations
Elected Officials Methodology
How PolicyLogic assesses governors, mayors, senators, and representatives on their campaign and inaugural commitments. Version 2.1 — operative standard. All prior versions superseded.
How assessments are produced. Assessments are produced by an automated pipeline and verified by researchers against the public record using documented procedures. Verification is ongoing rather than complete. Each assessment is labeled with its current state.

What Counts as a Promise

A statement qualifies as a trackable promise if it meets all three conditions: it is attributable directly to the official (not a surrogate) in a campaign or inaugural context; it is forward-looking (an expressed intention, not a description of existing policy); and it is verifiable — it has at least one condition that can in principle be confirmed true or false.

Statements of value without action components ("I believe in education") are logged but not assessed. They carry no verifiable condition, so they cannot be measured on delivery — but they are not discarded. Each is recorded on the assessment as a Clarity-1 Stated Value, marked Not Converted, and displayed alongside the assessed promises without a delivery record attached. This keeps the value-versus-record gap visible: an official does not bank rhetorical credit for a value simply because it was too vague to measure. Promises contingent on federal action outside the official's jurisdiction are excluded. Where a promise appears at multiple specificity levels, the more specific version is retained.

There is no cap on promise count. All qualifying promises are assessed. Capping promise count introduces editorial bias — selecting 6 from 30 means 24 are excluded on undisclosed grounds. Every assessment displays a mandatory disclosure: N promises tracked · estimated M identified · V stated values logged (not assessed) · X excluded for non-verifiability. Here estimated M is the reviewer's count of candidate promise-statements surfaced in the source sweep before qualification filtering; it is an estimate of the field, not an audited total, and is disclosed to show the share of statements that survived to assessment.

Promise Types

PolicyLogic recognizes three promise types, each with its own delivery scoring track. All three feed into the same D0–D4 delivery buckets — the distinction is in how the bucket is determined.

TypeDefinitionHow Delivery Is Measured
Quantitative Promise includes a numeric target or measurable threshold. Percentage of stated goal achieved. D4 = 70% or more, D3 = 40% to under 70%, D2 = 10% to under 40%, D1 = under 10% with documented action, D0 = no meaningful action.
Qualitative Promise specifies an action or outcome without a numeric target (pass a bill, make an appointment, launch a program). Action milestone ladder: D0 = no action · D1 = public commitment only · D2 = formal action initiated (bill introduced, passed committee) · D3 = advanced past the decisive hurdle (passed a full chamber, nominee confirmed, program launched) · D4 = fully delivered.
Negative Promise to prevent, avoid, or not do something ("I will not raise taxes"). Inverted criteria, assessed binary. D4 = condition fully avoided through the term. D0 = condition occurred. Intermediate buckets (D1–D3) do not apply to negative promises — a thing is either avoided or not. The single exception: any mid-term redefinition of the condition triggers an automatic Redefined flag and caps delivery at D2.

Assessment Framework

Each promise is assessed on three independent axes. The axes combine into a Promise Record. Promise Records are aggregated into a final Grade. Maximum per promise: 25 points.

Axis 1 — Delivery (max 12 points)

Delivery measures what the official actually accomplished relative to what was promised. It is the primary axis and carries the most weight in the final Grade.

CodeLabelDescriptionPts
D4Delivered70%+ of quantitative goal achieved, or qualitative promise fully delivered.12
D3Substantially Delivered40% to under 70% achieved, or qualitative promise advanced past the decisive hurdle (D3 on the ladder) with documented evidence.9
D2Partial10% to under 40% achieved, or qualitative promise at formal-action stage (D2 on the ladder). Concrete actions taken without full outcome.6
D1MinimalUnder 10% of a quantitative goal with documented action, or a qualitative promise at public-commitment-only stage (D1 on the ladder). No enacted policy or measurable outcome yet.3
D0Not DeliveredNo meaningful action taken, promise abandoned, or condition of a negative promise occurred.0

Their Role modifier (0.0–1.0): The official's direct causal contribution to the outcome is assessed and applied to Delivery points before aggregation.

Adjusted Delivery = Delivery Points × Their Role

See the Their Role lookup table below for anchor values by office-action combination.

Axis 2 — Difficulty (max 5 points, earned proportionally)

Difficulty rewards ambition — but only when delivery occurs. Zero delivery earns zero Difficulty points regardless of how ambitious the promise was.

CodeLabelDescriptionMax Pts
H3StructuralMulti-year initiative requiring coalition-building, constitutional change, or federal coordination.5
H2LegislativeRequires passage of legislation, significant political capital, or cross-chamber negotiation.3
H1ExecutiveAchievable through executive order, budget allocation, appointment, or regulatory change within direct authority.1
Difficulty Earned = Bucket Points × (Delivery Points / 12)
Difficulty Earned scales with unmodified Delivery Points, not Adjusted Delivery. Their Role discounts the Delivery axis only; it does not compound into the Difficulty axis. This is deliberate — an official who genuinely attempted an H3 initiative but was outvoted retains the ambition credit their effort earned, even where their causal share of the outcome was low.
Example: H3 promise with D2 delivery = 5 × (6/12) = 2.5 pts · H3 with D0 = 0 pts

Axis 3 — Impact (max 8 points)

Impact measures what was at stake. It is not modified by delivery — a failed promise on a critical issue still carries high impact stakes. This is intentional: if you promise something that matters and don't deliver, the Grade reflects how much it mattered.

ScaleDescriptionPtsMagnitudeDescriptionPts
S3Systemic / Statewide4M3Transformative4
S2Regional / Citywide2M2Significant2
S1Neighborhood / Segment1M1Minor / Symbolic1
Specificity Cap (Clarity Rule): If a promise has a Clarity rating of 2 (see Clarity Scale below), Magnitude is capped at M1. A vague promise cannot be awarded high magnitude because its intended scope was never defined.

Promise Record

Promise Record = Adjusted Delivery + Difficulty Earned + Impact
Maximum per promise: 12 (Delivery) + 5 (Difficulty) + 8 (Impact) = 25 points

Assessment Calculation

The final Grade is calculated from two ratios combined in a fixed 60/40 split. This split is published on every assessment and does not vary by official, party, or jurisdiction.

Delivery Ratio = Sum of Adjusted Delivery Points ÷ (N promises × 12)
Record Ratio = Sum of Promise Records ÷ (N promises × 25)
Assessment Input = (Delivery Ratio × 0.60) + (Record Ratio × 0.40)
The 60/40 weighting makes Adjusted Delivery the dominant term: it is the whole of the Delivery Ratio and the largest component of the Promise Record, so delivery drives roughly three-quarters of the Assessment Input. Difficulty and Impact adjust the result around that delivery core rather than substituting for it.
Design limit, disclosed: an official who makes only easy, low-impact promises and delivers them all can still place high. The framework records kept promises and does not penalize unambitious ones. Difficulty and Impact are surfaced separately on every assessment precisely so a reader can see what was promised, not only how much was delivered. A final Grade is not a substitute for reading the promise slate.
Assessment Input is a value from 0 to 1. It is multiplied by 100 and rounded to one decimal place before banding into the final Grade. All band thresholds are inclusive of their lower bound and exclusive of the next band's lower bound: B is 70.0 up to but not including 85.0, A is 85.0 up to but not including 90.0, A+ is 90.0 and above. The same inclusive-lower-bound convention applies to every range in this methodology — Delivery percentages, Difficulty bands, and Time Pressure cutoffs — so no value falls into two bands or into none.
A+90.0%+Exceptional delivery on ambitious, high-impact promises across multiple domains.
A85.0–89.9%Strong delivery across most domains with documented outcomes.
B70.0–84.9%Above-average to solid delivery; meaningful gaps remain.
C50.0–69.9%Mixed to marginal delivery; structural or political barriers evident.
D30.0–49.9%Poor to very poor delivery; few concrete outcomes.
FBelow 30.0%No meaningful delivery on tracked promises.

Worked Example

Worked Example · Housing Promise · Senator (Qualitative, H2)

"I will pass a tenant protection bill in the first session."

Promise TypeQualitative
Delivery BucketD2 — Bill introduced, passed committee, stalled on floor (formal action initiated)
Delivery Points6
Their Role0.6 — Advocated, dependent on legislature
Adjusted Delivery6 × 0.6 = 3.6 pts
DifficultyH2 (Legislative) = 3 max pts
Difficulty Earned3 × (6/12) = 1.5 pts
ImpactS2 + M2 = 2 + 2 = 4 pts
Promise Record3.6 + 1.5 + 4 = 9.1 / 25
Worked Example · Tax Pledge · Governor (Negative, H1)

"I will not raise the state income tax."

Promise TypeNegative
Delivery BucketD4 — no income-tax rate increase enacted or signed during the term (condition avoided)
Delivery Points12
Their Role0.8 — For negative promises, Their Role reflects whether the official's own action caused the avoidance: 1.0 where the official held and used veto or agenda control to prevent the increase; 0.4 or below where the increase failed for reasons unrelated to the official (e.g. the legislature never advanced one). Here: governor publicly committed and vetoed one rate-increase bill.
Adjusted Delivery12 × 0.8 = 9.6 pts
DifficultyH1 (within direct authority via veto) = 1 max pt
Difficulty Earned1 × (12/12) = 1.0 pt
ImpactS3 + M2 = 4 + 2 = 6 pts — unmodified by delivery; a kept fiscal pledge of statewide scope carries full stakes.
Promise Record9.6 + 1.0 + 6 = 16.6 / 25
Evidentiary StandardA D4 on a negative promise is established by the absence of a qualifying Tier 1 record across the term (no enacted rate increase in the session record or signed-legislation index), corroborated by one Tier 2 source confirming no such measure was enacted. Proving a non-event requires a complete negative search of the primary record, not a single document.

Their Role — Lookup Table

The Their Role modifier is the single most judgment-dependent element of the methodology. The following anchor values reduce inconsistency across AI runs and human reviewers.

ValueSituationCommon Examples
1.0 Sole or near-sole authority. Official acted unilaterally within clear constitutional or statutory power. Governor signing an executive order; mayor appointing a department head.
0.8 Official championed, negotiated, and signed legislation. Legislature was a necessary co-actor but official drove the outcome. Governor securing and signing a major budget deal; senator authoring and passing a bill with party in majority.
0.6 Official advocated consistently but was dependent on others who were not fully aligned. Senator in majority facing moderate opposition within own caucus; governor working with a split legislature.
0.4 Supporting or facilitative role. Outcome primarily driven by other actors, market forces, or federal policy. Mayor benefiting from a federal infrastructure grant they applied for but did not design.
0.2 Minimal causal influence. Outcome largely driven by forces entirely outside the official's sphere. Economic improvement during a governor's term primarily driven by national trends.
0.0 No meaningful causal connection. Official's actions were irrelevant to the outcome. Official claims credit for a federal policy they had no role in; outcome occurred despite official's opposition.

Clarity Scale

Clarity measures how specifically a promise was stated at the time it was made. It is assessed as of the original statement — it cannot be improved retroactively by subsequent clarifications.

Assessment begins at Clarity 2. Clarity 1 exists as a record-only tier: a pure values statement with no action component is logged and displayed but never enters the assessed set, so it receives no Delivery, Difficulty, or Impact points. It is marked Not Converted on the assessment. This is why a Clarity-1 statement can appear on an assessment yet carry no delivery record.

RatingLabelDefinition & Example
1Stated Value — Not ConvertedPure value or belief, no action component. Logged and displayed but not assessed; carries no verifiable condition. "I believe in stronger communities."
2Directional, no specificsGeneral intent stated, no mechanism, target, or timeframe. Magnitude capped at M1. "We will improve public safety."
3Specific policy namedA named policy, program, or bill is referenced. "I will pass a tenant protection bill."
4Specific + conditionsNamed policy plus timeframe, jurisdiction, or population. "I will pass a tenant protection bill in the first year."
5Specific + measurable targetNamed policy with a quantified, verifiable outcome. "I will reduce violent crime 20% by 2026."

Time Pressure Adjustment

Every promise is assigned a timeline bucket based on the official's stated or implied delivery window. Delivery is adjusted for whether a promise is early, on-track, or overdue. Reductions are a fixed step schedule, not continuous; the maximum penalty is 40%. Being early never reduces delivery — an official is not penalized for outpacing their own timeline.

ConditionStatusEffect
Time Pressure < 0.5EarlyAssessment is provisional and displayed with a clock indicator. No reduction applied; delivery is measured at full weight against progress to date.
Time Pressure 0.5 to under 1.0On TrackFull delivery weight. No adjustment.
Time Pressure 1.0 to under 1.25Overdue (mild)10% delivery reduction.
Time Pressure 1.25 to under 1.5Overdue (moderate)20% delivery reduction.
Time Pressure 1.5 to under 1.75Overdue (significant)30% delivery reduction.
Time Pressure 1.75 and aboveOverdue (severe)40% delivery reduction. Maximum reduction; does not increase further.

Behavioral Flags

Behavioral flags override or cap standard assessment when the official actively distorted the promise. Each flag requires a written rationale in the assessment JSON.

When more than one flag applies to the same promise, a cap (maximum) always takes precedence over a floor (minimum), and the lowest applicable cap controls. Example: if Externally Blocked sets a D2 floor and Scope Reduced caps Magnitude while Redefined caps delivery at D2, the delivery result resolves to D2 — the cap binds. A flag that sets delivery to a fixed value (Reversed = D0) overrides both caps and floors.

Reversed
Delivery = D0 regardless
Official explicitly reversed or repealed a previously promised policy.
Redefined
Outcome capped at D2
Definition of success materially changed mid-term, or condition of a negative promise redefined.
Externally Blocked
D2 minimum if actions taken
Promise failed due to documented external intervention meeting the Changed-Circumstances Test.
Credit Overclaimed
Their Role capped at 0.4
Official claimed credit for outcomes driven by prior administration or federal action.
Deadline Shifted
Time Pressure cap relief removed
Timeline extended without explanation after original deadline passed.
Scope Reduced
Magnitude capped at M1
Promise delivered in significantly diminished form without acknowledgment.

Transparency Flags

Transparency flags are applied at the assessment level and communicate limitations to readers without altering the result. They appear as labeled badges on the assessment.

Contested
A classification is disputed by the official or a credible third party. Evidence and rationale are documented.
Limited Evidence
Fewer sources than the standard minimum. Judgment applied. Assessment should not be cited as definitive.
Provisional
Produced by the automated pipeline; researcher verification against the public record is ongoing.
Under Review
Classification actively being re-evaluated due to new evidence.
Low Promise Count
Fewer than five promises assessed. Assessment may not reflect the full record.
Mid-Term Departure
Official left office before term ended. Rules for departure cases documented in full methodology.
Stated Value · Not Converted
A Clarity-1 values statement logged for the record but not assessed. Displayed without a delivery record so the value-versus-record gap stays visible.

Data Sources

Changelog

Version 2.1 — Clarifications, one substantive rule change, and a terminology standardization. No change to the axis structure, point maxima, or the 60/40 weighting.