Lenny's PodcastKeith Coleman & Jay Baxter: How bridging finds neutral truth
Through bridging-based scoring that rewards agreement between users who disagree; only 7% of proposed notes ever ship, and Meta now copies the algorithm.
CHAPTERS
- 0:00 – 6:56
What Community Notes is (and why it’s “context,” not traditional fact-checking)
Lenny kicks off with a clear explanation of Community Notes: users propose notes on misleading posts, other users rate them, and only notes that earn broad trust get shown. Jay emphasizes that the system’s goal is to add context so users can decide for themselves, not to act as a centralized fact-checking authority.
- •Users propose notes on any post they find misleading
- •Other contributors rate notes for helpfulness
- •Notes are shown when people who usually disagree find them helpful
- •Distinction: adds context vs. authoritative fact-checking
- •Examples like out-of-context/AI-generated images illustrate the need
- 6:56 – 12:06
Inside the “bridging-based” algorithm: agreement among past disagreers
Jay explains the core insight that makes the system work: it doesn’t rely on majority vote or instant publishing. Instead, it looks for surprising agreement between people who historically disagree, which tends to produce neutral, accurate notes and helps resist manipulation.
- •Not majority rules; not instant publishing
- •Algorithm seeks agreement across historically opposing raters
- •Surprising cross-group agreement predicts neutrality and accuracy
- •Bootstraps without external ‘ground truth’ labels
- •Anti-manipulation properties emerge from cross-group requirement
- 12:06 – 13:33
Eligibility, contributor onboarding, and why every post can be noted
The team lays out the philosophy that any post can receive a note—no exemptions for elites, politicians, advertisers, or Elon. Contributors must earn the ability to write notes by demonstrating good rating behavior, reinforcing quality and trust.
- •Every post is eligible for notes (including leaders and ads)
- •Contributors must earn writing privileges through rating quality
- •Random selection + basic criteria (e.g., verified phone)
- •System designed to feel fair, open, and trustworthy
- •Real-world example of correcting out-of-context imagery
- 13:33 – 17:25
Scale and impact: billions of views, near-million contributors, and smart matching
Keith shares surprising scale metrics: hundreds of notes shown per day, tens of billions of views annually, and a contributor base nearing one million. They also explain how notes can be matched across identical media/URLs to expand coverage while keeping notes specific and useful.
- •Hundreds of notes per day vs. ~10 traditional fact checks/day (comparison)
- •2024: ~95k notes seen ~30B times; rapid growth vs prior year
- •~950k contributors worldwide with more on waitlist
- •Media matching applies one note to many identical-image posts
- •Emphasis on specificity and “zero-click” context over generic warnings
- 17:25 – 21:31
Publishing thresholds and quality control: conservative by design
They unpack how notes are selected to show, including the notion of a publishing threshold (described as 0.4 on an internal scale) and additional filters for incorrectness. They also reveal that only ~8% of proposed notes get shown because protecting trust via high quality is the top priority.
- •Threshold is on an internal model scale (not a simple % of people)
- •Additional filters prevent ‘helpful but incorrect’ notes from showing
- •Team optimizes for trust: better to show fewer notes than a bad one
- •~7–11% of proposed notes are shown; ~8% typical
- •Bad-note behavior can cost users their ability to write notes
- 21:31 – 26:36
Polarization, reputation, and “all of humanity” participation
Lenny asks about extreme viewpoints and polarized users. Jay and Keith explain how the algorithm and reputation system handle them: ratings that consistently oppose cross-group consensus get down-weighted, but participation is intentionally broad to model what’s helpful to humanity overall.
- •Highly polarized topics may yield fewer publishable notes
- •External studies show notes change agreement with a post’s claims
- •Reputation system can stop counting consistently low-signal ratings
- •Philosophy: include all of humanity to learn what’s broadly helpful
- •Volunteers are intrinsically motivated; high-impact examples reinforce this
- 26:36 – 29:37
Behavioral effects: notes dramatically reduce resharing and can trigger deletions
Jay describes large engagement impacts: posts with notes see significant drops in likes and reposts, with network effects leading to roughly 50–60% reductions in resharing. They also discuss an unintended tradeoff: authors are more likely to delete posts after being noted, which can remove high-quality notes from view.
- •No ranking demotion needed; users organically reshare less
- •A/B test: ~30–40% drops in likes/reposts when notes appear
- •External research: ~50–60% drop in total reposts after a note
- •Notes ‘take the wind out’ of viral misinformation cascades
- •Authors become ~80% more likely to delete noted posts (tradeoff)
- 29:37 – 40:33
Origin story: Keith’s shift from management to solving misinformation at scale
Keith traces his motivations back to joining Twitter during the 2016 election and seeing information battles play out daily on the platform. After years in leadership and a moment of reflection (plus paternity leave), he proposes stepping away from managing PMs to pursue a high-impact experiment that became Birdwatch/Community Notes.
- •2016 election highlighted Twitter as a primary arena for public discourse
- •Existing approaches (fact checkers/internal decisions) didn’t scale or earn trust
- •Keith sought a ‘crazy idea’ with real-world impact
- •Thermal-style autonomy enabled exploration and prototyping
- •Early work began with research, prototypes, and pilots
- 40:33 – 46:01
Small teams, big impact: the Thermal project model and operating principles
They explain “Thermal”: a protected, high-autonomy model for building disruptive products inside a big company. Key ingredients include a single accountable ‘founder,’ a single senior decision-maker, full-time focus, and lightweight planning that supports fast iteration.
- •Thermal = isolation from bureaucracy + dedicated staffing
- •One clear project driver and one external senior decision-maker
- •100% focus increases iteration speed dramatically
- •Avoids heavy quarterly planning/OKR cadence; milestones drive goals
- •Early team composition: ~5 people (FE, BE, ML, design, research)
- 46:01 – 50:34
Algorithm evolution and internal competition: from PageRank-style to bridging
Jay details early algorithm attempts that prioritized anti-manipulation via PageRank-like methods, which didn’t solve bias when one side outnumbered another. Using pilot data, the team realized polarization was the core challenge and ran an internal ‘bake-off’ (Kaggle-style) to develop the bridging approach.
- •First production approach: anti-manipulation PageRank variant
- •Pilot data revealed bias/polarization as the main failure mode
- •Bridging-based agreement became the key requirement
- •Internal competition accelerated algorithm innovation
- •Inspired by prior work (e.g., cross-partisan likes, Polis) but adapted to notes
- 50:34 – 58:33
How the team operates day-to-day: one long doc, daily syncs, and minimal tooling
The conversation shifts to execution mechanics: the team coordinates via daily meetings and a long-running Google Doc rather than heavyweight task management. They argue that lightweight planning reduces time spent grooming backlogs and keeps priorities aligned with real-world urgency.
- •Daily team meetings to align on priorities and unblock launches
- •A long-running Google Doc serves as the coordination hub
- •Minimal Jira usage only for cross-team dependency requests
- •Avoids heavy backlog grooming that can crowd out real priorities
- •Goals and roadmaps change frequently based on observed problems
- 58:33 – 1:05:22
Working with Elon and the “opt-in” culture: lean teams, ownership, and deleting code
Keith shares what surprised him about X under Elon: extremely lean teams can move faster and feel more like owners. Jay adds the engineering discipline required to operate lean—deleting code and simplifying systems to reduce maintenance burden after major org changes.
- •Lean teams can ship ‘impossible’ scale-ups in weeks, not months
- •Opt-in principle: people self-select into a high-intensity environment
- •Shrinking bureaucracy increases pace of launches and experimentation
- •Cross-team collaboration improves when everyone feels like an owner
- •Engineering lesson: deleting/auditing code can matter more than adding features
- 1:05:22 – 1:10:44
Launching Birdwatch carefully: low expectations, pilots, and sifting ‘gold’ from noise
Keith describes disciplined rollouts that required the product to prove itself at each step, from prototypes to MTurk-like tests to a small public pilot. They debated signaling risk (even considering a dumpster fire GIF) and focused on identifying high-quality notes reliably.
- •Stepwise validation: mockups → user research → pilot → expansion
- •MTurk-style testing showed laypeople could produce ‘gold’ notes
- •Public pilot began with ~1,000 contributors to manage risk
- •Main challenge became identifying the best notes, not writing them
- •Early public attention even reached Elon before acquisition
- 1:10:44 – 1:18:12
Core principles: people-driven, no override button, and radical transparency
Keith lays out foundational principles that made Community Notes credible: it must represent the people’s voice, not the company’s, and there can be no internal ‘kill switch’ for notes. Transparency is equally central—open-source code and public data enable audits, replication, and external trust-building despite real engineering costs.
- •Voice of the people, not the platform’s editorial stance
- •No internal button to change a note’s status once it qualifies
- •Problems must be fixed at the system level, not via ad hoc moderation
- •Open-source code + public datasets allow full replication and auditing
- •Engineering tradeoffs were made to ensure external runnability and verification
- 1:18:12 – 1:26:12
Stress test in global crisis: Israel–Hamas misinformation surge and speed improvements
They recount a major real-world stress test: the early days of the Israel–Hamas war produced an overwhelming volume of misleading imagery and claims. Community Notes handled it with hundreds of notes in days, aided by recent launches like media notes, matching, and faster note publication times.
- •Conflict triggered massive misinformation volume and rapid virality
- •~500 notes in the first few days covering out-of-context and fake media
- •New capabilities (image/video notes + matching) proved crucial
- •Recent speed-up reduced time-to-note; median ~5 hours during surge
- •Polarized topics still produced cross-group agreement on verifiable facts
- 1:26:12 – 1:32:14
Anonymity/pseudonymity: counterintuitive lever for honesty and cross-partisan agreement
Addressing an audience question, Keith explains why the team shifted to pseudonymous contributors. Anonymity reduced fear of harassment and increased willingness to cross partisan lines, improving both participation and the honesty needed for bridging-based consensus.
- •Real-handle attribution reduced participation on controversial topics
- •Pseudonymity reduces harassment risk and increases contribution volume
- •Anonymity increases cross-partisan willingness to rate honestly
- •Quality remains high due to multiple system safeguards
- •Parallel insight: private likes can similarly encourage honest behavior
- 1:32:14 – 1:47:57
Surviving leadership changes and the road ahead: AI-assisted notes and ‘SuperNotes’
They discuss how Community Notes survived multiple CEOs by proving value with data and shipping consistently through upheaval. Looking forward, they aim for “more, better notes faster,” including new features like improved note requests, and explore AI-human collaboration (SuperNotes) that could even allow the public to meaningfully shape the core algorithm.
- •Longevity came from product value + consistent execution amid chaos
- •Cost savings wasn’t the driver; trust, speed, and scale were
- •Near-term: improve ‘request a note’ (bat signal) and core algorithms
- •AI frontier: LLMs generate variants + simulated juries to predict helpfulness
- •Vision: product and even scoring improvements increasingly built by the public