← Back to projects

Midterm Messaging 2026

A live tracker of what about 300 House, Senate and governor candidates say in their own campaign messaging, collected daily and labeled by topic, tone and target through Election Day 2026.

Tool·In progress, live through Nov 3·September 2026

What are 2026 House, Senate and governor candidates actually talking about? This project collects about 300 candidates' own campaign messaging every day from campaign websites, YouTube uploads, Bluesky and X. It labels each item by topic, tone and target, then compares groups: party against party, favored against underdog, incumbent against challenger, and one channel against another. It runs as a live tracker through Election Day, November 3, followed by a post-election analysis.

Measurement

  • The candidate is the unit. Messages are labeled one by one, then rolled up into each candidate's topic shares, so one prolific poster can't outweigh thirty quiet ones.
  • "Sentiment" split into parts that mean something for campaign text: tone (promote, contrast or attack), stance toward the people and parties a message names, and how blame is framed within a topic.
  • Standing from prediction markets: market win probabilities sort candidates into favored, toss-up and trailing.
  • Comparisons stay within a channel. Each platform reaches a different audience, and nearly every candidate on Bluesky is a Democrat, so Bluesky is shown on its own and never pooled into party comparisons.

Pipeline

  • A scheduled daily Python job refreshes the race roster, collects new items from each channel, and rebuilds a DuckDB corpus (20,000+ items by October 1, growing by several hundred a day).
  • Messaging is joined to public context data: forecasters' race ratings, 2024 presidential results on the 2026 district lines (The Downballot), and Census ACS profiles of each district and state.
  • Collection is deliberately polite: an identifying user agent, robots.txt respected, per-host rate limits, and no workarounds for bot walls or paywalls. A blocked source is logged as a documented gap.
  • Collectors are idempotent and can backfill, so a missed day costs nothing but that day's website snapshots.
  • A hand-verified roster of 147 races and 298 candidates is the source of truth for IDs and links.
  • Snapshots publish to this site's Postgres database through a least-privilege database role, and the page regenerates on demand.

Labeling and validation

  • A 45-topic codebook with written boundary rules, frozen before the full labeling pass.
  • About 150 hand-labeled items, with the test set labeled blind (no model labels visible).
  • Several model setups and a keyword baseline were compared on a tuning set; the finalists were scored once on the blind test set, with accuracy broken out by party, since party-specific jargon can make a classifier's errors lopsided.
  • Run-to-run label noise was measured, and each item is labeled once and stored.

Status

Started in September 2026 and live through Election Day. The tracker now shows which party talks about what (overall, week by week, race by race, and for any two groups of candidates), what candidates put on their websites against what they post and run as ads, topics one party almost never raises, lines that travel word for word between candidates, each party's distinctive words, how alike opponents' topic mixes are, what campaigns change on their websites week to week, and whether messaging follows constituents' circumstances. Every finding opens to the candidates behind it, with short excerpts linked to the originals. A post-election analysis follows.