No PriorsRethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein
CHAPTERS
- 0:00 – 0:59
AI agents as a new security threat: non-human actors with real permissions
The episode opens with a warning: threats are no longer just humans and ransomware, but AI agents acting as non-human identities inside enterprise environments. Because these agents can have legitimate access, the risk is less about “breaking in” and more about rapid, high-velocity misuse or mistakes. The hosts frame the core challenge as maintaining visibility, control, and recovery capabilities as agent activity accelerates.
- •Shift from human-only threats to agent-driven threats with legitimate permissions
- •Velocity and scale of damage increases (e.g., rapid table drops, destructive actions)
- •Non-technical employees can deploy agents without understanding compliance/security
- •Agents can chain (agents using other agents), complicating accountability
- •Need to assume breach and design for detection + recovery
- 0:59 – 1:50
Meet Eon: cloud backup and disaster recovery reimagined for the AI era
Elad introduces Eon and its founders, positioning the company as cloud backup/disaster recovery built for modern AI needs. The conversation sets up a broader exploration of how data infrastructure must adapt when AI systems need historical, contextual data—and when security risks intensify. Eon’s premise: data protection and data usability are converging.
- •Eon is positioned as backup/disaster recovery designed for AI-era requirements
- •Conversation will connect data infrastructure, AI usage, and enterprise change
- •Motivation: managing and using long-lived enterprise data becomes strategic
- •Shift from pure protection to also enabling AI workflows
- 1:50 – 2:40
Eon’s “data foundation”: discover, classify, ingest, and query across clouds
Gonen explains Eon’s core product: a cloud-based data foundation that maps and classifies data across environments and hyperscalers. Eon then ingests structured and unstructured sources into a cost-effective store that supports protection/recovery plus search, query, and AI model application. The goal is to make scattered data both secure and usable.
- •Map and classify data across multiple hyperscalers and environments
- •Identify sensitive vs. non-sensitive data and where it lives
- •Ingest structured and unstructured data into a unified foundation
- •Enable querying/search plus applying LLMs/AI on top of the ingested data
- •Deliver cost-efficient storage while supporting protection and recovery
- 2:40 – 3:33
From backups to business value: full-history data becomes an AI asset
Elad highlights the insight that backup repositories contain long-term customer history and operational context that can unlock new applications. Ofir argues that AI tailwinds make data the most durable advantage for companies because models and compute are increasingly commoditized. This reframes backup/DR as strategic infrastructure for extracting value, not just insurance.
- •Backups capture full historical context that typical pipelines may omit
- •AI increases the value of long-lived, proprietary enterprise data
- •Models/compute are more ephemeral; data becomes the durable differentiator
- •Backup + recovery infrastructure can double as an AI-ready data layer
- 3:33 – 6:12
Data as moat: why Google bought Spirit Airlines’ data (not planes)
Ofir and Elad discuss the Spirit Airlines bankruptcy dataset purchase as a signal that enterprise datasets are becoming prime strategic assets. They note multiple AI-adjacent bidders and predict more data acquisitions, including outreach to companies asking to buy proprietary datasets. The broader point: proprietary data is increasingly treated as “gold” and a competitive moat.
- •Google reportedly bought Spirit’s data for ~$10M, not physical assets
- •Multiple bidders indicate rising market value for enterprise datasets
- •Startups and labs increasingly ask companies to sell their data
- •Old archival data ("tapes on a shelf") is becoming valuable again
- •Data plus people become key competitive advantage as tooling commoditizes
- 6:12 – 9:48
Training agents needs real-world data: scarcity, synthetic limits, and masking
They explore why high-quality datasets are hard to find and why real enterprise data helps agents behave realistically. Spirit’s data could be useful not just for airline workflows but as a rare snapshot of a large enterprise’s internal operations and hierarchy. They also emphasize the tension between usefulness and sensitivity (PII/financial data), pushing toward masking and controlled access.
- •Agents require high-quality, realistic training/post-training data
- •Public datasets are limited; many look synthetic or lack real-world texture
- •Enterprise datasets capture operational reality (hierarchies, processes)
- •Need to manage sensitive content (PII, financials) via masking/controls
- •Trend: combine novel synthetic approaches with selective use of real data
- 9:48 – 11:05
Why existing data tooling breaks down: data is locked behind org silos and risk
Ofir explains that traditional data work was tactical: teams found known data, built pipelines, and delivered predefined outputs. In the AI era, leadership demands broad activation of all enterprise data—but ownership is fragmented across business units, legacy systems are poorly understood, and pulling data risks uptime, compliance, and security. This creates incentive conflicts that block AI initiatives.
- •Old model: data teams worked tactically on known datasets and projects
- •AI pressure from CEOs/boards demands broad enterprise data activation
- •Business units often don’t know what data they have or where it resides
- •Legacy systems and “mystery servers” complicate extraction and governance
- •Extraction can threaten production uptime and trigger compliance/security issues
- 11:05 – 15:34
Eon’s approach to unlocking data safely: semantic context, continuous ingest, least risk
Eon positions itself as tooling for the data leader to discover and understand enterprise data without negotiating every silo manually. By classifying data, building context/semantic layers, and continuously ingesting without impacting production, Eon aims to enable AI use while preventing accidental leakage of sensitive fields. The emphasis is on controlled sharing, cost efficiency, and operational safety.
- •Help data leaders find, map, and classify all organizational data
- •Add context via semantic layers to make data usable for AI workflows
- •Continuously ingest relevant data without compromising production systems
- •Prevent accidental sharing of sensitive information via classification
- •Store/access data efficiently for both governance and AI activation
- 15:34 – 18:14
Security in the agent era: ransomware lessons, detection signals, and rapid recovery
Gonen recounts an AWS-era ransomware incident where incomplete tagging meant large portions of an environment weren’t actually protected—helping motivate Eon. The same detection concepts (irregular writes, entropy changes) apply to agent-caused incidents, but agent actions occur far faster. The chapter underscores resilience: detect anomalies early and recover granularly and quickly.
- •Real-world ransomware exposure can stem from missing tagging/classification
- •Detection methods: irregular write patterns, entropy changes, anomaly signals
- •Core resilience goals: protect, detect, and recover granularly
- •Agent-driven incidents mirror ransomware patterns but happen much faster
- •Customers increasingly fear internal misuse by approved agents
- 18:14 – 22:11
How agents reshape the enterprise stack: non-human identity, endpoints, and governance
Ofir argues agents will increase the need for dashboards to understand complex chains of action and responsibility. Non-human identity becomes a top security problem as agents activate other agents across systems, including endpoints like laptops connected to internal and external networks. Meanwhile, boards push rapid AI enablement, but CIOs/CISOs face rising fear of uncontrolled data exposure by non-technical “builders.”
- •Dashboards may increase as humans need observability into agent activity
- •Non-human identity management becomes a major cybersecurity category
- •Endpoint risk returns as agents run on laptops bridging multiple networks/apps
- •Board/CEO push to enable AI conflicts with security/governance concerns
- •Non-technical staff can deploy tools that mishandle sensitive company data
- 22:11 – 27:00
Re-imagining data infrastructure: from single-purpose pipelines to context-rich activation
They argue existing ETL/plumbing was designed for narrow, predefined questions, so context across systems wasn’t critical. With AI, the value shifts to collecting, cleaning, and connecting data across the organization so teams can ask richer questions and build new capabilities. The growth of agent-generated data creates both opportunity and noise—requiring new tools for discovery, cleaning, and usability.
- •Traditional pipelines reflect single-purpose intent and limited cross-context
- •AI rewards unified context: combining datasets unlocks new insights/use cases
- •Organizations ingest far more data than before; growth is “out of proportions”
- •Agent-generated data adds both value and significant noise
- •Need tools to locate, clean, and make distributed data usable across sources
- 27:00 – 27:52
Cost and efficiency in the AI era: storage sprawl, token spend, and “value per token”
As data scales and becomes more distributed, accessing and processing it becomes expensive—not just storage but inference and token costs. They note the shift away from “token maxing” toward ensuring each token spend produces measurable value. This reinforces the need for efficient storage formats and smarter data access patterns.
- •Data sprawl creates both access challenges and cost pressure
- •AI workloads add token/inference costs on top of storage and compute
- •Shift from experimentation to disciplined ROI: “value for every token”
- •Motivation for efficient storage and selective, well-governed activation
- 27:52 – 33:04
Cloud shift vs. AI shift: faster transformations, loss of control, and new go-to-market motions
Drawing on their CloudEndure/AWS migration experience, they compare cloud migration’s heavy, human-driven effort to AI’s faster, more chaotic transformation. AI adoption is accelerated because it’s easier to understand (ChatGPT moment) and driven by existential pressure from leadership, but it also increases fear around breakage, leakage, and IP exposure. They highlight emerging deployment patterns like forward-deployed engineers, faster sales cycles, PLG-to-enterprise motion, and even acquisition-driven AI modernization.
- •Cloud migrations were massive but slower and labor-intensive
- •AI transformations are faster and can cause customers to “lose control”
- •AI is more universally understood than cloud, increasing top-down pressure
- •New motions: forward-deployed engineers accelerate legacy enterprise adoption
- •Trends: faster PLG adoption, shortened cycles, and M&A-led AI modernization
- 33:04 – 34:50
Wrap-up: the AI/data transition is just beginning
They close by emphasizing that most companies are still early in AI adoption despite intense awareness and pressure. Data is broadly recognized as critical, but modernization is both necessary and scary. The episode ends with thanks and standard show outro.
- •Most companies are still at the beginning of AI adoption
- •Organizations recognize data importance but struggle with processes and risk
- •Modernization is unavoidable despite governance and security fears
- •Final reflections and closing thanks