Skip to content
YC Root AccessYC Root Access

Diamond Maps: Efficient Reward Alignment for Generative Models

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Douglas Chen about Diamond Maps, a method for steering generative models toward desired outputs more efficiently. Reward alignment depends on estimating how promising an intermediate state in the generation process is. Existing flow-map methods make this estimate using a single possible final output. Diamond Maps instead samples multiple outcomes from the same intermediate state, producing a better estimate and stronger guidance. The work includes both a fine-tuning method and a training-free inference-time method, allowing existing generative models to be aligned without necessarily retraining them. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostDouglas Chenguest
Aug 6, 20265mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
August 6, 2026
Duration
5m
Channel
YC Root Access
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Douglas Chen about Diamond Maps, a method for steering generative models toward desired outputs more efficiently. Reward alignment depends on estimating how promising an intermediate state in the generation process is. Existing flow-map methods make this estimate using a single possible final output. Diamond Maps instead samples multiple outcomes from the same intermediate state, producing a better estimate and stronger guidance. The work includes both a fine-tuning method and a training-free inference-time method, allowing existing generative models to be aligned without necessarily retraining them. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

SPEAKERS

  • Ankit Gupta

    host

    Host/interviewer on YC Root Access (Y Combinator).

  • Douglas Chen

    guest

    Researcher presenting “Diamond Maps,” a method for efficient reward alignment of generative models.

EPISODE SUMMARY

In this episode of YC Root Access, featuring Ankit Gupta and Douglas Chen, Diamond Maps: Efficient Reward Alignment for Generative Models explores diamond Maps enables efficient reward alignment via stochastic value estimation Diamond Maps targets reward alignment: steering a strong base generative model toward a user-defined task or style via reward guidance.

RELATED EPISODES

Evaluating the Fine-Grained Planning Abilities of Web Agents

Evaluating the Fine-Grained Planning Abilities of Web Agents

ChartNet: Training Vision-Language Models to Understand Charts

ChartNet: Training Vision-Language Models to Understand Charts

Any-Horizon Reasoning for Video Agents

Any-Horizon Reasoning for Video Agents

LeanAgent: Lifelong Learning for Formal Theorem Proving

LeanAgent: Lifelong Learning for Formal Theorem Proving

Zero-Shot Predictive Models for Relational Databases

Zero-Shot Predictive Models for Relational Databases

Improving Small Language Model Reasoning With A* Search

Improving Small Language Model Reasoning With A* Search

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.