CHAPTERS
- 0:00 – 2:06
LLMs as “ceiling raisers”: solving problems humans can’t
Mikhail frames LLMs as more than productivity boosters: they unlock solutions that were previously out of reach, even with unlimited human effort. He describes a “centaur” workflow where the best results come from iterative back-and-forth between human and model.
- •LLMs enable entirely new capabilities, not just faster execution
- •Example of a long-standing technical problem unlocked with model assistance
- •Most valuable work is collaborative: neither human nor model alone can solve it
- •Raising the ceiling is more impactful than raising the floor
- 2:06 – 2:43
From floor to ceiling: why ROI gets harder to quantify
Boris contrasts easy-to-measure “floor raising” (time saved) with the harder question of valuing “ceiling raising” (new things now possible). The conversation shifts from individual wins to organizational measurement and executive communication.
- •Floor gains are measurable via time saved and tasks automated
- •Ceiling gains create new product/business options, harder to benchmark
- •Leadership needs a framework to reason about capability expansion
- •Sets up Shopify’s approach to measuring engineering productivity
- 2:43 – 3:42
Shopify’s measurement stack: PRs, project complexity, and throughput
Mikhail explains Shopify’s unusually rigorous approach to quantifying productivity improvements. They normalize output by project complexity and track execution speed through internal systems to estimate percentage gains.
- •Beyond PR count: measuring complexity and project scope
- •Using internal tracking (GSD system) to evaluate throughput
- •Normalizing team output by complexity to avoid misleading metrics
- •Quantifying productivity gains as percentages to guide decisions
- 3:42 – 4:12
CTO guidance: default to the largest model (and learn what you’re missing)
He argues that large models reveal “unknown unknowns,” so leadership should start with the best model to see what’s possible. This philosophy drives Shopify’s decision to provide broad access and heavy model usage.
- •Largest models maximize ceiling-raising potential
- •You can’t evaluate what smaller models miss without comparison
- •Start with maximum capability, then optimize after evidence
- •Shopify’s stance: unlimited tokens and heavy model deployment
- 4:12 – 4:44
Why Shopify gives engineers unlimited tokens
Mikhail connects token abundance to experimentation and capability discovery. He emphasizes investing in both inference/training capacity and optimization to keep the organization moving fast.
- •Unlimited tokens encourage exploration and reduce friction
- •Running “heaviest models” to push capability boundaries
- •Investment in GPU inference/training plus optimization
- •Organizational learning accelerates when access isn’t constrained
- 4:44 – 5:46
Digital twins for merchants: predicting growth from sequences of actions
Mikhail describes Shopify’s HSTU prediction system: modeling a company as a sequence of actions, analogous to language modeling. This enables building a “digital twin” to simulate outcomes and guide interventions that increase merchant success.
- •Company behavior modeled as sequential actions (like tokens in text)
- •Digital twin enables simulation without risking real-world harm
- •Counterfactual experiments: change shipping speed, ads, credit, loans
- •Outputs translate into concrete merchant interventions in production
- 5:46 – 6:45
From simulation to production impact: interventions that move the bottom line
The digital twin isn’t just research—Shopify uses it operationally to decide which actions to recommend or trigger. Mikhail ties this directly to merchant outcomes and financial results.
- •System identifies sequences most likely to drive merchant growth
- •Examples: offer loans, push operational recommendations, launch campaigns
- •Safe experimentation enables broader, faster optimization
- •Clear connection to Shopify’s financial and business outcomes
- 6:45 – 7:28
“Previously it would take infinity”: defining ceiling-raising in practice
Asked how long the system would have taken before, Mikhail argues it effectively wasn’t feasible. The key value is entering a solution space that wouldn’t have been considered at all.
- •Talent concentration, foresight, and data work made it impractical before
- •LLMs change feasibility, not just timeline
- •Ceiling-raising = projects that weren’t in the consideration set
- •A new category of strategic options becomes available
- 7:28 – 8:13
AI for engineering leadership: detecting slippage before it happens
Mikhail shares a major mindset change: LLMs can help manage teams, not just code. He built a system that analyzes signals and flags likely deadline risks early enough to intervene.
- •He initially believed LLMs wouldn’t help management—then reversed
- •System monitors ongoing work and predicts schedule risk
- •Proactive alerts enable earlier issue resolution
- •Improves overall organizational functioning and execution reliability
- 8:13 – 9:56
How the EM role changes: supervising engineers who supervise LLMs
Mikhail argues technical leadership becomes more important as LLMs accelerate building. Because models can “please” and amplify poorly framed asks, managers need stronger judgment and better task dispatch across humans and models.
- •Simple execution becomes easier; technical acumen matters more
- •Risk: models mirror the question—bad asks become fast bad outputs
- •Managers shift toward orchestrating work across multiple LLM agents
- •Focus moves from status-checking to decision quality and coordination
- 9:56 – 10:57
Advice to CTOs: avoid token limits, add circuit breakers, use a barbell strategy
He recommends not constraining tokens too aggressively, while still preventing runaway spend. His “barbell” approach: use the biggest models for development/research, and optimize production with right-sized or fine-tuned models.
- •Don’t over-limit tokens; you can’t see the counterfactual capability
- •Use circuit breakers for runaway processes and misuse
- •Barbell: largest model for coding/dev/testing/research
- •Production: fine-tune and right-size models for cost/throughput/quality
- 10:57 – 11:28
Common mistake: penny-pinching on dev, overspending in production
Mikhail criticizes organizations that restrict development tokens but run overly large models in production. He frames correct allocation as a competitive advantage: invest where learning and leverage are highest, economize where scale costs dominate.
- •Biggest leverage is in development and exploration, not production runtime
- •Fine-tuned/cheaper production models free budget for frontier dev usage
- •Misallocation pattern: saving on dev while wasting in production
- •He expects to outcompete teams not following this strategy
- 11:28 – 12:18
Closing anecdote: Claude, Linear A, and racing the future
The conversation ends with a humorous story about Mikhail’s personal ambition to decipher Linear A—until someone did something similar using Claude. It reinforces the episode’s theme: capabilities are advancing so fast that waiting means missing the window.
- •Personal “retirement project” displaced by rapid model-enabled progress
- •Illustrates the speed of ceiling-raising breakthroughs
- •Highlights the urgency of experimentation and adoption
- •Wrap-up and farewell
