Skip to content
No PriorsNo Priors

No Priors Ep. 91 | With Cohere Co-Founder and CEO Aidan Gomez

In this episode of No Priors, Sarah is joined by Aidan Gomez, cofounder and CEO of Cohere. Aidan reflects on his journey to co-authoring the groundbreaking 2017 paper, “Attention is All You Need,” during his internship, and shares his motivations for building Cohere, which delivers AI-powered language models and solutions for businesses. The discussion explores the current state of enterprise AI adoption and Aidan’s advice for companies navigating the build vs. buy decision for AI tools. They also examine the drivers behind the flattening of model improvements and discuss where large language models (LLMs) fall short for predictive tasks. The conversation explores what the market has yet to account for in the rapidly evolving AI ecosystem, as well as Aidan’s personal perspectives on AGI—what it might look like and when it could arrive. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @AidanGomez Show Notes: 0:00 Introduction 0:36 Co-authoring “Attention is all you need” 2:27 Leaving Google and founding Cohere 4:04 Cohere’s mission and models 6:15 Pitfalls of current AI 8:14 How enterprises are deploying AI today 10:58 Build vs. buy strategy for AI tools 14:37 Barriers to enterprise adoption 20:04 Which types of companies should pretrain models? 24:25 Addressing flaws in open-source models 25:12 Current and expected progress in scaling laws 29:54 Advances in multi-step problem solving and reasoning 32:29 Key drivers behind the flattening curve of model improvements 36:25 Exploring AGI 39:59 Limitations of LLMs 42:10 What the market has mispriced

Sarah GuohostAidan GomezguestElad Gilhost
Nov 21, 202444mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:27

    Aidan Gomez’s path into AI: Toronto, Hinton, and a “lucky” Google Brain internship

    Sarah opens by situating Aidan’s background—growing up in Canada and landing in the center of modern deep learning. Aidan attributes his trajectory to being immersed in the University of Toronto AI ecosystem and a serendipitous Google Brain internship that set everything in motion.

    • University of Toronto’s AI culture and Geoff Hinton’s influence
    • Early immersion in deep learning as a default career aspiration
    • Google Brain internship with Lukasz Kaiser as a turning point
    • Aidan’s anecdote about accidentally getting a PhD-track internship as an undergrad
  2. 2:27 – 4:03

    From Transformers to Cohere: seeing the GPT-2 trajectory and deciding to build

    Aidan describes moving across major research hubs and collaborators from the Transformer era to Pathways-scale training efforts. GPT-2’s release clarified the direction of “web-scale” language models, motivating him and his co-founders to start Cohere to build and apply the tech.

    • Work across Google Brain (Transformer team), U of T, and European research collaborators
    • Pathways as an early vision of massive training infrastructure
    • GPT-2 as a signal of the technology’s trajectory
    • Calling co-founders to pursue building large language models outside Google
  3. 4:03 – 4:40

    Cohere’s enterprise-first mission: not a ChatGPT clone, a deployment platform

    Aidan frames Cohere’s purpose as enabling other organizations—especially enterprises—to adopt LLMs to improve productivity and transform services. The company explicitly prioritizes enterprise needs over building a consumer chatbot competitor.

    • Enterprise adoption as Cohere’s core value proposition
    • Focus on workforce productivity and product transformation
    • Positioning against building a direct ChatGPT competitor
    • Platform + products to operationalize LLMs for businesses
  4. 4:40 – 6:14

    Models plus reliability: why enterprise AI is foundation + go-to-market + product

    Cohere balances research/model development with enterprise requirements like security, reliability, and customer support. Aidan highlights a growing product emphasis aimed at shortening time-to-value by helping customers avoid recurring implementation mistakes.

    • Models are necessary but insufficient for enterprise success
    • Enterprise requirements: reliability, security, support
    • Shift toward productization to reduce implementation friction
    • Learning from repeated customer POC failures across the market
  5. 6:14 – 8:13

    The biggest pitfall: brittle prompting and RAG details that make or break deployments

    Aidan points to overestimating models as “human-like” and underestimating how sensitive they are to formatting and prompt structure—especially in RAG systems. Cohere’s response is twofold: make models more robust and wrap them in more structured APIs and workflows.

    • LLMs are highly sensitive to prompt/data presentation quirks
    • RAG failures often come from retrieval formatting, storage, and system details
    • 2023 saw many enterprise POCs fail due to unfamiliarity and idiosyncrasies
    • Two fixes: increase model robustness and add structured product/API guidance
  6. 8:13 – 10:58

    What enterprises build today: Q&A, summarization, and workflow acceleration

    Cohere sees broad enterprise demand centered on document Q&A and summarization, with large productivity gains despite “mundane” surface descriptions. Aidan gives concrete examples in manufacturing and healthcare, where models can compress massive document or record review into seconds.

    • Common use cases: enterprise knowledge Q&A and summarization
    • Manufacturing example: manuals/diagnostics chat for frontline workers
    • Healthcare example: summarizing decades-long patient records for clinician briefings
    • Impact comes from speed and coverage humans can’t match at appointment time
  7. 10:58 – 14:19

    Build vs. buy in AI apps: the pyramid strategy and competitive-advantage projects

    Sarah asks about equilibrium between specialist vendors and in-house builds; Aidan predicts a hybrid “pyramid.” Organizations should buy commodity/general tools, but build bespoke capabilities that map to their unique advantage—illustrated via an insurance RFP response assistant.

    • Pyramid model: general copilots at the base, bespoke domain tools at the top
    • Hybrid future: buy what’s standardized; build what’s uniquely strategic
    • Insurance example: RAG-based research assistant accelerating RFP responses
    • Horizontal LLMs are like CPUs; customer insight determines winning applications
  8. 14:19 – 16:01

    Barriers to enterprise adoption: trust, security/privacy, and know-how

    Aidan argues the primary blocker is trust, especially around sensitive data in regulated industries. Cohere differentiates by flexible deployment (on-prem/VPC) and emphasizes that the second barrier—skills and familiarity—will ease over a few years.

    • Trust as the core adoption barrier
    • Security/privacy constraints in finance and healthcare; data can’t leave locked environments
    • Cohere’s deployment flexibility (on-prem, in/out of VPC) to reach sensitive data
    • Developer knowledge gap as a time-bound barrier (2–3 years to permeate)
  9. 16:01 – 17:53

    Is there a trough of disillusionment? Why adoption work will take years regardless

    Aidan acknowledges some hype-cycle dynamics but argues the tech is still improving and unlocking new applications frequently. Even without new models, he believes there is “half a decade” of integration work to embed current capabilities into real systems and workflows.

    • Not yet in a classic trough; still early with steady capability gains
    • New application categories keep opening every few months
    • Even with no further model advances, years of integration/product work remain
    • Shift from debating usefulness to executing broad deployment in the economy
  10. 17:53 – 20:00

    Specialization playbook: fine-tune first, then post-train, and pre-train only when necessary

    Sarah probes how enterprises should specialize models; Aidan presents a cost-gradient approach. Start with the cheapest lever (fine-tuning), move to post-training as needed, and only touch pre-training for requirements like new languages or large proprietary corpora.

    • Different needs require different levers (tone/format vs language competence)
    • Example: Japanese model with Fujitsu requires pre-training intervention
    • Recommended progression: fine-tuning → post-training (SFT/LHF) → limited continuation pre-training
    • Aim to touch only a small fraction of pre-training when possible
  11. 20:00 – 21:50

    Who should pre-train models: continuation runs, enterprise data moats, and ‘empirically wrong’ takes

    Aidan disputes the idea that only AGI labs should pre-train, arguing large enterprises with massive token stores can benefit. He emphasizes continuation pre-training as a practical middle ground—smaller runs that meaningfully adapt models without starting from scratch.

    • Pre-training doesn’t make sense for SMBs/startups, but can for data-rich enterprises
    • Continuation pre-training can be ~$5M rather than $50M+ full runs
    • Enterprises can leverage proprietary data at scale as a meaningful advantage
    • Direct rebuttal: claiming no one outside frontier labs should pre-train is ‘empirically wrong’
  12. 21:50 – 24:16

    Cohere’s model strategy amid open source: spend enough, build cheaper, and target enterprise needs

    Aidan explains Cohere’s bar: there’s a minimum viable spend to build truly useful models, but costs fall over time. Cohere aims to be ~6–12 months behind frontier models, deliver what enterprises need at sensible price points, and maintain capital efficiency even with supercomputer costs.

    • Minimum spend threshold to achieve enterprise-useful model quality
    • Training can be far cheaper than first movers (e.g., GPT-4-class enterprise needs for $10–20M)
    • Strategy: don’t lead the frontier burn; be slightly behind but better aligned to customers
    • Supercomputer costs are high, but the business can still be profitable
  13. 24:16 – 25:09

    Why not just use open-source models: gradients, levers, and vertical integration

    Sarah challenges the need to train in-house given open-source options like LLaMA. Aidan argues that frozen, post-trained open-source releases limit the levers available for fine-tuning and data intervention, whereas vertical integration gives Cohere more control to meet enterprise requirements.

    • Open-source releases arrive ‘cooled down’ with no gradient access in training pipeline
    • Fine-tuning alone can be less effective than controlling data and training process
    • Vertical integration provides more levers (data, training choices, adaptation depth)
    • Control translates into better ability to serve specific enterprise constraints
  14. 25:09 – 32:25

    Scaling laws and the new ‘reasoning’ unlock: inference-time compute as a product lever

    Aidan says scaling gains are flattening and “vibe checks” are no longer sufficient; expert evaluation in domains matters more. He highlights reasoning models as a major unlock, shifting improvement economics from only capex-heavy retraining to on-demand intelligence via more inference-time compute.

    • Progress is flatter; distinguishing generations requires domain experts and targeted evals
    • For many enterprise tasks, models are already close to ‘good enough’ with customization
    • Reasoning enables multi-step internal deliberation and self-correction
    • New lever: pay for more inference-time compute to get smarter outputs immediately
  15. 32:25 – 36:23

    Why progress is flattening: fine-detail data bottlenecks and the limits of human knowledge

    Using an oil painting analogy, Aidan explains that early gains were broad strokes, while frontier improvements require fine, scarce data. Synthetic data works well where answers are verifiable (code/math), but biology/chemistry require expert distillation, eventually hitting the boundary of what humans know—implying experiments as a longer-term path.

    • Early progress used broad internet data; frontier gains require fine-grained targeted data
    • Synthetic data helps in verifiable domains (code, math)
    • Real-world sciences face expert-data bottlenecks and eventual scarcity
    • Long-term path may require models to run experiments, but that’s hard to scale soon
  16. 36:23 – 44:15

    AGI, takeoff, and LLM limits: continuous progress, practical friction, and market mispricing

    Aidan frames AGI as continuous rather than a binary milestone, rejecting superintelligence takeoff narratives due to practical constraints and real-world friction. He argues transformers are general but not always efficient, and closes with a market view: models aren’t truly commoditized—price dumping masks a small number of capable providers during a long “tech refactor.”

    • AGI as a continuum; skepticism toward ‘God-like’ takeoff and self-improvement fantasies
    • Practical constraints and friction limit naive extrapolation of curves
    • Transformers can model many sequences; inefficiency exists vs domain-specific structures
    • Market mispricing: ‘commoditization’ is often loss-leading; few providers + massive rebuild implies pricing pressure later

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.