Skip to content
YC Root AccessYC Root Access

ChartNet: Training Vision-Language Models to Understand Charts

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with MIT PhD candidate Jovana Kondic about ChartNet, an open-source data generation pipeline and million-scale dataset for chart understanding. Charts require models to combine visual recognition, text understanding, and numerical reasoning. ChartNet generates diverse examples by translating charts into plotting code, augmenting that code, and rendering new images with corresponding tables, summaries, and reasoning traces. Training on ChartNet improved open-source models across a range of chart tasks and transferred to real-world benchmarks, showing how carefully structured synthetic data can give smaller models capabilities commonly associated with much larger systems. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostJovana Kondicguest
Aug 6, 20265mWatch on YouTube ↗

CHAPTERS

  1. 0:07 – 0:28

    ChartNet overview: a synthetic, open-source dataset for chart understanding

    Jovana introduces ChartNet as a foundation-model data generation pipeline and dataset designed to improve robust chart understanding in vision-language models (VLMs). She frames it as a large, diverse, fully synthetic resource created with open access in mind.

    • ChartNet targets next-generation VLM chart understanding
    • Fully synthetic and open source dataset + generation pipeline
    • Built to meet foundation-model scale and diversity demands
    • Collaboration between MIT and IBM research groups
  2. 0:28 – 0:58

    Why charts are hard for VLMs: text parsing + numerical reasoning

    The conversation clarifies what “charts” include and why they differ from natural images. Models must read embedded text and perform numerical/relational reasoning over plotted values and axes.

    • Charts require understanding both visual elements and text labels
    • Numerical reasoning is central (values, trends, comparisons)
    • Covers simple to complex charts (bar, scatter, violin, etc.)
    • Charts resemble real scientific figures (e.g., Nature paper plots)
  3. 0:58 – 1:11

    Design goals: comprehensive chart diversity plus rich multimodal metadata

    Jovana describes the dataset’s ambition: the largest and most diverse to date, with broad chart-type and library coverage. Beyond images, each sample includes supporting artifacts that help models learn deeper structure and semantics.

    • Million-scale ambition and broad diversity
    • Includes image plus code, data tables, and natural language summaries
    • Multimodal coverage to support different learning signals
    • Built to match foundation-model training needs (size + modalities)
  4. 1:11 – 1:42

    Key insight: charts are programmatically generated

    ChartNet leverages the fact that many charts originate from plotting code. This allows the pipeline to reconstruct and manipulate the underlying code to produce controlled variations and additional annotations.

    • Charts typically come from plotting scripts
    • Code provides a structured representation of the chart
    • Enables systematic generation and augmentation
    • Supports producing aligned image-code pairs
  5. 1:42 – 2:19

    Two-stage pipeline: image-to-code seeding, then LLM-driven code augmentation

    Jovana outlines the core pipeline: start with a seed set of chart images, use a VLM to infer approximate plotting code, then iteratively augment that code using an LLM. This yields scalable, diverse synthetic examples along with metadata derived from the code and data.

    • Stage 1: VLM translates seed chart images into approximate plotting code
    • Stage 2: LLM performs iterative code augmentation for diversity
    • Code can generate additional metadata (tables, summaries, etc.)
    • Produces aligned training tuples grounded in executable code
  6. 2:19 – 2:36

    Iterating in code space: rendering variations from a base chart

    Ankit restates the approach and Jovana confirms the workflow with a concrete example: start from a simple line chart, generate code, then produce multiple augmented variants by modifying code and re-rendering. The visible images are outputs of controlled programmatic changes.

    • Start with a simple chart and infer code
    • Iterative augmentations happen directly in code space
    • Charts are re-rendered by executing the modified code
    • Rendered images help visualize what’s changing under the hood
  7. 2:36 – 2:54

    Quality filtering: keeping synthetic data reliable

    Jovana notes a quality filtering component that enforces standards on the synthetic outputs. This step is positioned as essential to ensuring the dataset remains useful for training robust models.

    • Dedicated quality filtering unit
    • Ensures synthetic samples meet predefined quality criteria
    • Helps prevent noisy or invalid chart-code pairs
    • Supports dependable large-scale training data
  8. 2:54 – 3:24

    What each sample contains: code, image, tables, summaries, and Q&A reasoning traces

    The dataset structure is described as a rich tuple per sample. It includes the plotting code and rendered image, plus derived data tables, natural language summaries, and question-answer traces that include reasoning steps for supervision.

    • Each example includes plotting code + chart image
    • Includes data table and natural-language summary
    • Includes Q&A traces with reasoning
    • Also covers multiple chart types and plotting libraries
  9. 3:24 – 3:39

    Additional subsets: real-world curated data, human annotations, and grounding

    Beyond the core synthetic collection, ChartNet includes extra slices of data aimed at realism and evaluation coverage. These subsets include curated real-world charts, human-annotated samples, and grounding-focused data.

    • Core synthetic dataset is supplemented by additional subsets
    • Real-world curated samples included
    • Human-annotated data provided
    • Grounding-oriented data to connect visuals to underlying values
  10. 3:39 – 3:55

    Fine-tuning results: consistent gains across model families and tasks

    Jovana reports experiments fine-tuning open-source models from sub-1B up to 7B parameters. She observes consistent improvements on chart-understanding tasks, with gains transferring to real-world public benchmarks despite synthetic training data.

    • Fine-tuned models from <1B to 7B parameters
    • Improvements observed across multiple chart understanding tasks
    • Gains generalize to public real-world benchmarks
    • Synthetic training data can still improve real-world performance
  11. 3:55 – 4:05

    Efficiency highlight: small open models outperform larger frontier models on chart tasks

    A standout claim is that a ~2B parameter open-source model can outperform much larger GPT-4o on chart understanding after ChartNet fine-tuning. The takeaway is that targeted, multimodal dataset design can unlock capabilities often attributed only to massive models.

    • ~2B open-source model surpasses GPT-4o on chart tasks (as reported)
    • Dataset quality and structure can substitute for sheer scale
    • Multimodal supervision (image + code + text + values) is key
    • Positions ChartNet as a recipe for capability-efficient training
  12. 4:05 – 5:15

    Industry adoption: IBM uses ChartNet in VLM training mixture; strong benchmark performance

    Jovana shares that IBM Research used ChartNet as part of its core training mixture for a next-generation VLM. She cites results where a ChartNet-trained model outperformed leading open models and Claude Opus on chart/table extraction benchmarks.

    • ChartNet incorporated into IBM Research VLM training mixture
    • Strong performance on chart and table extraction benchmarks
    • Outperforms leading peer open-source models (as reported)
    • Beats Claude Opus on dedicated tasks with fewer parameters
  13. 5:15 – 5:39

    Access and community: open-source release, Hugging Face availability, ongoing curation

    The conversation closes with how to access ChartNet (paper + dataset links/QR codes) and early adoption metrics. Jovana emphasizes ongoing expansion and invites community feedback on improvements and future directions.

    • Fully open source; accessible via links/QR codes (and referenced Hugging Face)
    • Reported ~50,000 downloads since release
    • Actively curated and expanding over time
    • Call for community feedback on next steps

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.