Skip to content
Aakash GuptaAakash Gupta

Mikhail Episode 3

Aakash Gupta and Mikhail on building an AI-native product org using agents, knowledge graphs, rituals.

Aakash GuptahostMikhailguest
Jul 25, 20261h 6mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:001:30

    Intro

    1. AG

      What's gonna happen to my job?

    2. MI

      What actually will happen is that teams will be smaller, leaner, and faster. Not only it's gonna change, the lines are gonna blur.

    3. AG

      Meet Mikhail Shcheglov, the CPO at OLX Classifieds, formerly CPO at Azerbaijan's biggest e-commerce marketplace and GPM at Bolt. This is, like, the coolest visualization I've ever seen. So what tool is this?

    4. MI

      It understands our industry, classifieds, our market, our business model specifically at the level of 54%.

    5. AG

      What's your process these days to hire PMs? How do you find PMs that are at this level of AI native?

    6. MI

      I typically look at three things: fundamental problem-solving, systematic thinking.

    7. AG

      On top of our own job, we need to keep the pipeline of really good PMs. How can this help with recruiting?

    8. MI

      This automates already, like, 70, 75% of the recruiting workflow.

    9. AG

      What's gonna happen to the future? Is product management going to exist in the future?

    10. MI

      I have good news for you. I think-

    11. AG

      Before we get into today's show, please take a second to check that you're subscribed on YouTube and following on Apple and Spotify Podcasts. If you want access to all of my favorite AI tools, I've gotten them to give you an entire year of their paid plans. Check out bundle.aakashg.com for an entire year of Bolt.new, Airtable, Speechify, Descript, Magic Patterns, Linear, Dovetail, Arise, and Mobbin. And now, into today's show. I

  2. 1:302:37

    Why AI-native teams need a “company operating system” (and what it replaces)

    1. AG

      know you guys are overwhelmed with guides on Hermes and OpenClaw and Claude Code and ChatGPT. It seems like there are a million AI tools out there, and more and more of you are reporting to me that you're not even feeling more productive after using AI. So I've been searching for those PMs and product leaders who are not just getting, like, a 10% boost in productivity, but are getting a two, three X boost in productivity, and Mikhail is one of those people. What he has built at OLX Classifieds is an entire operating system. If you've been hearing about buzzwords like OpenClaw and Hermes and saying there's no value for PMs, this episode is the one that will change your mind. He has got his entire team hooked into this operating system that he has built, and they are validating features, shipping features, reviewing features. The possibilities for a PM are endless, and I can't wait for him to share this all with you. Mikhail, welcome to the podcast.

    2. MI

      Glad to be here.

    3. AG

      So how do people really build a company operating system? How do they measure it? What are the keys to getting the most out of AI as a company?

  3. 2:373:38

    Digitizing organizational knowledge to prevent context leakage

    1. MI

      All right, so, um, if you think about this, like, what's the core goal of having AI native teams is to be able to automate as much knowledge as we can. 'Cause, like, in, in previous, I call this old paradigm. Like, you used to have a knowledge worker, like an employee who has built a context around a certain domain. Let's say it's, it's a deep expertise domain like search, recommendations, ML, and stuff like that. And then all of a sudden, this person decides to leave the company and takes all the knowledge with him or with her. Um, he's a bottleneck, effectively, a knowledge bottleneck. And the core thing that we're trying to prevent here is the leakage of that knowledge. Because if you could have a single storage of this entire business, customer, product, technical knowledge in one place, then, uh, it would increase, like, the entire value of your organization. That's just one angle to it. And another angle

  4. 3:385:39

    The OLX knowledge graph: layers, signals, and what “54% coverage” means

    1. MI

      is the better, uh, your AI knows your context, the more autonomy you can give it, the higher level tasks you can delegate to it. And, um, let me explain how it actually works in our case. So what you see here is, uh, the entire knowledge graph of our company that has been formed over, like, five months. Like, five months is basically nothing, and you can already see this, like, web of interconnected particles. And these particles, they cover everything, uh, from our contacts, who talks to whom, to, uh, which projects we're working on, to which customers we've talking with, which businesses we've reached out, even our funnel metrics, basically everything. And you can see in the top right corner, like, there is a specific metric which is a product context coverage, and this is one of my personal KPIs. Because the higher this percentage of coverage, the more, as I mentioned, you can delegate to it. Because right now, um, what, what, what this means is that it understands our industry, uh, like classifieds, our market, our business, our business model specifically, the top line, the bottom line, the drivers behind it, our customers and the value drivers for them at the level of 54%, which already means that it can operate as a capable, I would say, junior to mid product manager and even make backlog level decisions. Um, the more this, uh, coverage grows, like at numbers of 70 to maybe, like, 90%, I would not be surprised if this knowledge graph would enable you to do some strategy level work and help business navigate decisions. It's a question of when we're gonna get there, but we'll definitely gonna get there. Let me maybe, um, zoom

  5. 5:398:00

    Using the graph to detect silos and evaluate team discovery strength

    1. MI

      in a little bit on, on the, um, uh, elements of this knowledge graph. So, um, you have three layers. So the first layer is the product, so all the nodes that are connected to the product side of things. And you can see, like, product has been, uh, very involved into everything. Um, you have the contacts, which is basically all the people talking to each other in different parts of the organization Uh, and there are, like, personal things that, uh, I share, like my thoughts, my reflections, um, and the personal things that some of the other people share, like, of course, fully anonymized. Then another layer is the teams, right? I, I don't really go at the, a level of actual, uh, like, say, f- cross-functional teams. It's more of a clusters or divisions, right? So if you take buyer and seller experience, you can immediately see how interconnected they are within the organization, or, like, the platform teams or, like, the pay and ship teams. Uh, and you can even drill down at the level of a specific product manager, and you can see on a team level what's their knowledge of their existing context, which for me is a very strong predictor of how good the team is in terms of customer discovery. Um, not only that, uh, there is another angle to it. If the team isn't really well-connected, and you can clearly see that this team has, like, like, very little, uh, nodes, overlapping nodes with, uh, the buyer and seller experience, even though a platform is a horizontal layer, and they should be- ideally, they should be intertwined, like, with each other. You can clearly see that, um, there is some, there some instances of silos happening, which means that teams should talk more with each other. And two, I would also raise the question of the stakeholder management. Like, are stakeholders really being involved in the processes within the teams, and so on? Um, and, um, that's kind of a, my core, um, tool that I'm using because it, it shows me the quality, uh, of an entire, uh, of the output of the entire organization. And, uh-

    2. AG

      This is, like-

    3. MI

      Yeah

    4. AG

      ... the coolest visualization I've ever seen. So what tool is this? How are these metrics working? How are they measuring things?

  6. 8:0012:36

    How the context-coverage metric is measured (and why it works)

    1. MI

      Um, so you can use, like, tools like Obsidian for this. I personally created this one from scratch using Fable because it was, like, a two-liner prompt, basically. Um, the, uh, important question that you asked is, um, like, super important one, which is how do we measure this? Um, and we have a very specific prompt that asks an AI, um, what percentage of, um, knowledge around industry, and you have to be very specific what do you imply by industry. In our case, we have five verticals. We have real estate, we have outdoor, we have good sales, we have services and jobs. Uh, what percentage of knowledge around the business, uh, meaning the business model, the actual PNL, the drivers, and what percentage around customers, like different buyer, seller, customer segments, even on a cohort level, marketing, and so on. Which percentage of knowledge do you have at the current point in time, taking all of your memories loaded into you directly? And what we've seen is that, uh, AI has been quite accurate in digitizing this abstract request into a specific number as an output. And what I've also, uh, what I'm also observing is that the more active product managers are, and they typically come with, I would say, transcripts or research studies or, like, RFD documents, and they load it directly into our agent, into its memory. And I'm seeing that the metric is actually improving over time. It doesn't mean that it's a single source of truth, and I think each organization should come up with a specific context metric that would be highly specific and relevant to them. Um, but this is something that I found directionally works.

    2. AG

      Here's a quick word from our sponsors. I need to say something that most AI coding companies don't wanna hear. Their tools are not built for enterprise. Think about what happens when 20 people in your org start building with AI. Marketing makes a dashboard, product prototypes a feature, sales builds a demo tool. Each one looks different. Different components, different styles, no shared design system, and none of the code is pullable. Engineering can't integrate any of it. You've got 20 AI-powered prototypes and zero production-grade output. Bolt.new solved this with something called design system agents. You upload your actual design system, your NPM packages, your CSS files, your component library. The DS agent builds a custom system from your real code. Now every person building in Bolt.new across your entire org is using your components, your buttons, your tokens, your brand. And when engineering pulls the code, it matches what they already ship with. That's the difference between an AI tool and an AI platform. Check it out at bolt.new/aakash. I used to think I had a retention problem. Turns out I had a messaging problem. I was sending the same onboarding emails to every new user, whether they activated on day one or never logged in again. I had no idea who was slipping or why. Customer.io changed that. Every message I send is now based on what users actually do in the product. Someone hits a key activation moment, they get nudged to the next one. Someone goes quiet, they get a different path entirely. Their AI agent makes it fast. I describe the campaign I want, and it builds the full journey for me. Triggers, timing, copy, even branching logic. And when I want to know how something is performing, I just ask the agent directly, and it tells me what to do next. They also have an MCP server, which means AI tools like Claude can see directly what's happening in your Customer.io workspace. Your segments, your customer data, your attribution, all of it. So instead of explaining your business context every time you need help, Claude already knows it. Notion used Customer.io to personalize their onboarding and hit nearly 50% open rate, improved conversion by six to seven percent with localized campaigns, and pushed open rates up another 20% through AB testing. The idea is simple. Customer.io helps you deliver more impact from every message you send. If you're a PM or founder and your onboarding is still one size fits all, try Customer.io at customer.io.

    3. MI

      For us.

    4. AG

      Okay, so this is a totally different way to view a product team. Most product teams, most PMs aren't used to this. To-- If they wanna build this, it sounds like there's two things they do. They define the metric, and then they- Create the visualization with Fable. Talk-- Let's take a step back here. What is the role of a CPO in an organization like this? What is the new agentic CPO doing here?

  7. 12:3615:18

    The “agentic CPO”: shifting from process scaffolding to AI operating systems

    1. MI

      Oh, that's, that's a great question, Aakash. I think, like, what has changed significantly is that in the past, um, like the whole team was responsible, uh, for the outcomes and outputs, right? Outcomes were measured in, uh, how much your metric has moved, and outputs, how much RFD documents, how much code, how much prototypes, uh, Figma mockups, like your team has been able to produce depending on the function. Uh, right now, what is, what has changed, and changed pretty significantly, is that the cost of this output is, like, negligible. Um, but what becomes more important is the quality and, uh, the token consumption. So what do I mean by quality? Quality for me is, um, the amount of, uh, AI outputs that have significantly moved or moved the metric. So it's, it's essentially outcomes, but outcomes that were driven by AI outputs, and this is how I measure it in terms of the metrics. In terms of the actual tooling, I think previously what, um, CPO was responsible for is, uh, building a scaff- you know, procedural scaffolding on top of the organization. What it means is that you have certain like hiring principles, right? Uh, your vision. Then, uh, if you go, uh, a level lower, you have certain processes like planning cadence, like a year- a yearly strategy, then a quarterly strategy, then a sprint planning. And then you have like different reviews, like design reviews, product reviews and so on. And, uh, this allowed the organization to function. But right now, what is more important, it's not the process itself, it's creating an operating system within which AI could be a collaborator of product managers, and where AI could not only help to, um, um, kind of, uh, come up or challenge ideas, but it could autonomously make decisions. Um, and the best way to achieve this is that, is by owning the entire agentic scaffolding architecture, uh, and creating specific rituals, uh, creating specific processes, and enabling the team to be in touch with those agents all the time, training the team to use it actually. Uh, let me know if you'd like to, to go, uh, deeper into any aspects of this.

    2. AG

      How does this manifest in the PM process? Where are the savings had for teams? Why should a CPO suddenly be spending so much time as an agentic orchestrator?

  8. 15:1816:56

    Where PM time is saved: automating rituals so PMs can focus on discovery

    1. MI

      The, the biggest reason for this, like if you break, if you break down the time, uh, that product managers on average, uh, used to spend, probably, uh, not, not probably actually, I estimated this based on my previous experiences, around 50% of the time of a typical product manager was spent, was spent on processes and rituals. Um, you have different, uh, business reports, weekly reports, stakeholder management, uh, reports, um, different demos and so on. And this is a, a repetitive type of task that, uh, doesn't really require much cognitive input in many cases. It's just something that you have to do manually. Uh, and if we abstract all of this and start delegating this part of work to AI, then essentially what you end up with is you have a product manager that is solely focused on discovery, like the, where the biggest leverage actually is, and a single product manager could actually do the work of two product managers, um, because he, he doesn't have the burden, uh, of, uh, of all the processes that a typical corporate environment imposes on him.

    2. AG

      Wow, okay. So we've seen the knowledge graph. We understand the role of the CPO and what it does for a PM. Can you show us what this agent looks like in practice, how a stakeholder might interact with it?

  9. 16:5618:51

    Slack agent in action: status reporting, workspace integration, and rule libraries

    1. MI

      There are multiple ca- buckets of tools that this, uh, agent can solve, and we constantly improve it. So the first bucket is, um, uh, status reporting. Like you can ask it like what's the status of a specific project? Let's say, uh, what's, uh, the status of auto spare parts catalog integration? Uh, respond in English because he might get finicky. Okay.

    2. AG

      [laughs]

    3. MI

      And while we're waiting for this, um, um, there are multiple other things. So status can be, you can get a, a response in whichever shape or form you want, um, directly in Slack, in Google Doc, uh, in Confluence, whichever format is more convenient for you. Second thing is, uh, it is integrated into all of the, uh, workspaces, right? Google Workspace, meaning Calendar, Gmail and so on. I don't read my email anymore, like an agent does it for me and pings me in case if it's, uh, if it's urgent. I don't manage my calendar anymore. And, uh, it's, it's es- it's especially difficult when, when we live in this interconnected world where everybody's remote in different parts of the globe and, uh, just setting a single meeting. Oh, you can clearly see it's all there. Scope listings automated. Yeah, you can clearly see like what's the status. Uh, it's quite straightforward.

    4. AG

      Nice. And did you like, have you been tweaking the prompts on how it gives status updates? 'Cause this is like a pretty good one.

    5. MI

      Yeah, yeah. So it has a massive, uh, library of rules and imperatives, I think like 700 lines or so, uh, to give you, uh, the cleanest and the most factual output without hallucinations.

    6. AG

      Very cool. So we'll go into how people build that. What other use cases should they know about within Slack that they can accomplish here?

  10. 18:5121:17

    Agent as gatekeeper: handling stakeholder feature requests before PM escalation

    1. MI

      Um, like o-one angle to this, let's say you're a stakeholder, right? And you have a certain feature, uh, that you just assume is important to have by what-whatever reason. Like a typical route for a stakeholder in this case is to go directly to a, a product manager and harass a product manager with feature requests, right? Uh, but now, um, like we have a gatekeeper, which is, uh, our agent running under the hood, and we train our stakeholders to interact with an agent before reaching out to a product manager.

    2. AG

      Oh.

    3. MI

      And let's say, yeah, I just woke up and I, I have a feature request. I urgently want to have, uh, a video submission, um, for autos, uh, story like in Instagram, uh, 'cause it's fancy.

    4. AG

      [laughs] 'Cause I saw somebody else doing it, right?

    5. MI

      Yeah.

    6. AG

      That's always the reason.

    7. MI

      All competitors are doing it, we are behind. Um, yeah, uh, for, for sellers, yes. Uh, and the-- this is just like a typical feature request which doesn't have a problem framing, it doesn't have any impact estimates, uh, it doesn't have any rationale behind, uh, and it doesn't respect the actual, uh, priorities that have already been defined for, for the organization. And, um, by doubt, we'll just give it a little bit of time to think. Um, it will respond with, uh, a list of clarifying questions. And in case if, uh, after answering all the clarifying questions, an agent decides that this feature isn't worth building, then it will politely say so. Uh, if it is worth building, then this feature will be added to the backlog and escalated to the responsible product manager. Uh, an agent has all the organizational structure ma-mapped out internally, so he knows exactly who the owner of a specific domain is.

    8. AG

      Wow. And do you consider yourself the owner of this Saul Goodman agent?

    9. MI

      Yes.

    10. AG

      A lot of people, they wanna outsource that to like an AI ops role or something like that.

  11. 21:1722:22

    Why the CPO should own the agent (not AI Ops): iteration speed and business impact

    1. MI

      Yeah. I, I think it's a great point, and it's a dangerous way to go. I think, um, uh, you should own this for two reasons. Like one reason is the speed of iteration. Like if you own this and you're constantly improving, and I'm getting feedback every single day, what works, what doesn't work, and so on, and I immediately like open my IDE or my Claude code and I make the changes right off the bat as we go. It's super fast from feedback to deployment. Uh, secondly, um, this actually has an impact on the organization because, uh, one, it saves time, uh, and two, it stirs decision-making, uh, which means it might have an impact on the entire business. And I think, uh, uh, CPO should own this. If you delegate this to an engineering team, or if you delegate it to, to another person who, who doesn't have the skin in the game, then you're not gonna get the same level of speed and impact.

    2. AG

      Okay. So CPOs should be building this. Now they wanna know how. Can we open up the covers and show people how they build something like this?

  12. 22:2225:25

    Under the hood: architecture choices (OpenClau + Hermes) and memory stack design

    1. MI

      Absolutely. Absolutely. Let me open up my IDE. Actually, I will show you two ways. Um, yep. This is just to give an overview. I will, I will kind of walk you through the architecture, uh, the brain, the memory, the tools, the skills, and I'll explain it all as we go. Uh, so for the architecture here, what we're using is, um, a combination of OpenClau and, uh, Hermes. Uh, why those two? Because OpenClau is just a great scaffolding. It has everything available out of the box and great, uh, engineering support. Um, but Hermes has a unique feature to it, which is an automated skill generation. And I've actually tested it, uh, across five core topics that, uh, my team works, and I've seen an improvement of, uh, in, in a recall metric, which means that it was way more accurate in responding using those skills by 31%, which was a dramatic improvement. So, uh, I decided to blend both of those scaffoldings together, uh, in a combination. Um, and it's, it's jacked in terms of the capabilities because we've been improving it for like five months continuously. Uh, when it comes to memory, it has three layers of memory. So the first layer is what you've seen is a knowledge graph. Um, it's like an entire universe of know- of interconnected knowledge and everything, and you can see the code here and so on. Um, but what is plugged in there, the, it's a second layer, which is a vector database. And the vector database is essentially every single piece of knowledge that an agent receives gets immediately converted into vector. And you need this because most of the requests are fuzzy. Um, and in order for you to get good retrieval, you, you do need to have vectors because you might ask some random stuff, right? That is, that is not related to any specific keyword, and an agent needs to be able to match your request to a specific, uh, point, to a specific, um, uh, number in your knowledge graph. And the third layer And what I've seen makes, uh, agents the most robust and they don't lose context is every single conversation is stored in transcripts. You can clearly see that every single day, uh, the agent stores everything into MD files. And not only this, uh, encompasses Granola transcripts, because all of our meetings are transcribed, but also all of the conversations that I had with an agent and all of the reflection that an agent has. So those three layers of memories, persistent memories, uh, they create this robust knowledge base which improves every single day.

    2. AG

      Wow, how does that exactly-- how do you architect that properly? So obviously you get a Granola enterprise license, you have that coming in and-

    3. MI

      Yep

  13. 25:2531:41

    Counterintuitive lesson: summarization hurts retrieval—store raw transcripts instead

    1. AG

      ... recording every single meeting. But I guess you don't wanna just save all the meeting details. You wanna strip out some of the personal conversation out of meetings, and then you wanna write these more condensed, synthesized MD files. How do you go from raw transcripts that might have personal information to useful synthesized output?

    2. MI

      That, that's a great point as well and, um, that was my assumption, that you actually have to summarize and synthesize the output to be useful. But we tested this, um, in, in, in multiple ways, and it turns out that summarization actually hurts a retrieval.

    3. AG

      Oh.

    4. MI

      Yeah. For, for two reasons. Like, one is when you summarize something, you lose granular details, and the devil is always in the nuance and the details. And another reason is that when you summarize, you impose a certain template, uh, onto whichever conversation that you had. For instance, uh, like the tasks that were solved, the tasks that remain, what is important, what is not important. And then on top of this, on top of this template, you, you, you try to kind of shove your, your transcript into this template. And, uh, we, we've noticed like a huge fidelity loss and, uh, I think it was like 20, 25% worse recall. That's why, uh, my personal architectural decision here was let's just store every single transcript, e- every single conversation that we had in the raw form, uh, without any summarization because it costs nothing.

    5. AG

      Quick thought experiment for you. Is there anything in this video you should be trying on your own? If there is, try it, take a screenshot, post it on LinkedIn or X and tag me. I'd love to see what you're learning. Now, a quick word from our sponsors before we get into the back half of the pod. Do you know how to take an AI product from idea to development to evaluation to deployment and eventually to scale? That's exactly what Product Faculty's AI PM certification helps you do. I even took the course myself. You'll learn directly from Rohan Varma, the product lead working on Codex at OpenAI. You'll go deep into AI prototyping, evaluations, agents, AI native workflows, Claude Code, OpenClau, latency, cost, guardrails, [chuckles] RAG, routing, fine-tuning, and production systems. You'll even build your own AI product as your capstone with unlimited one-on-one support. So if you wanna stop just learning AI and actually build AI products that work, join Product Faculty's AI PM certification on Maven. Five thousand plus students have graduated, and they have 1,000 plus reviews. Use code AAKASH550 to get $550 off your enrollment. I used to live in report purgatory. Every team had a different number. Every weekly review started with someone reconciling spreadsheets. We stopped hiring more analysts and gave the reconciliation to an AI employee instead. Victor is an AI employee that lives in Slack and Microsoft Teams. It connects to 3,000 plus tools your team already uses, ships real deliverables, and every action goes through your team for approval first. It's the closest thing I've seen to a small team running like a much larger one. Let me share three things that Victor does that changed how my team operates. First, ask Victor has replaced our Monday metric scramble. Someone types, "Flag any customer whose usage dropped 40% o- week over week, and draft an outreach loom for the account owner to approve." Ninety seconds later, the next step is ready. Second, scheduled tasks run the work that nobody wants to remember. Every morning, Victor checks overnight support tickets, drafts replies for the on-call to approve, and escalates anything mentioning churn. Nobody had to ask for it. Finally, Spaces ship internal tools in minutes. Ask for a renewals dashboard, Victor builds it. With auth in a database, and post the link to your channel. The team can stop opening four tabs to get the same view. So stop chatting with AI and start working with it. Get started at victor.com. There's $100 in free credits with no card required at V-I-C-T-O-R.com/AakashGupte3. You can find that link in the description. I want to take a second to talk to you about the fourth cohort of LAN PM Job. I trained 30 students in cohort one, 50 students in cohort two, and 75 students in cohort three, and I am bringing back the program for cohort four. It starts in August, and it lasts three months, where you're gonna have intense sessions, a Monday morning session where I go over your resume, behavioral interviews, LinkedIn. On top of that, Bart Jaworski is gonna be teaching you the PM fundamentals in 2026, how to write AI PRDs, how to AI prototype with Claude Code, all of the key skills you need to freshen up your knowledge for this market. And Ankit Vermani is gonna be teaching you AI product management. He is an AI product manager at Uber, and he is gonna teach you how to build AI features that actually work successfully. On top of that, Prasad Reddy is gonna be doing one-on-ones with you for mock interviews, LinkedIn review, candidate market fit review. So it is a full package. It is three courses in one for one low fee. So join at LANPMJob.com.

    6. MI

      Nothing.

    7. AG

      Wow. And Hermes and OpenClau on their own can figure out how to get into which meeting and which meetings are relevant context, and they won't just fill up their context window with random meetings?

    8. MI

      Yeah, absolutely. Because, like, what happens is that when you ask a request, what it does, it, it, it does a search query. Um, and it, it, it's a blend. It's a hybrid search query. Uh, it tries to do, uh, exact keyword matching. Uh, if it's not successful at this, uh, which I think 75% it's not because it's hard, uh, it's, it's a very ambiguous raw context, uh, then it does the, uh, vector search. And vector search allows you to, to do this exactly fuzzy retrieval based on probabilities, and that pretty much handles, like, the, the remainder 75% of use cases. So it only rece- uh, retrieves the relevant bit, pieces of, of data that is highly specific to your request, and it doesn't really overload the prompt with, uh, with, uh, tokens there.

    9. AG

      Very cool. So that's the memory component. What else do people need to know to build this?

  14. 31:4137:46

    Imperatives, tool integrations, and how Mikhail edits the system on the go

    1. MI

      I think the, the thing, uh, that after, after, uh, deciding on the architectural aspects of this, it is hugely important to create a list of, uh, imperatives. And, um, as you know, um, um, LLMs are very biased because of how they are trained, and, um, more specifically, um, they are focused on the resemblance of good output rather than the actual results.

    2. AG

      [chuckles]

    3. MI

      You know? And, uh, um, you can, you can, um, do workarounds, uh, to solve this, and the best way is to have this, uh, list of imperatives. For instance, uh, what, what, what for me is super critical is that there's no fabrications, right? Uh, there is, um... If you look here, um, it's, it's just a huge list of different imperatives that we've created over time. It's who you are. It's your voice. Uh, it's think before you act because that's, that's a terrible thing I've noticed, uh, LLMs do a lot. Uh, they provide you the output that looks plausible, but there wasn't that much thinking behind it because it's full of contradictions. Uh, then there's always facts over guesswork and anti-patterns. Another angle which, which super pissed me off a lot, I call this, uh, fake helpful. For instance, if you ask your agent to book you a meeting, and suddenly your tokens in the Google workspace has expired, and then the agent says, "Well, I'm, I'm sorry, I cannot do this," but it, it's super easy for you to do. Just open up a calendar, type in Google Calendar, uh, name your meeting, choose your time, choose per- I mean, it's useless. I mean, it's, it's an obvious advice. Uh, you don't even have to waste tokens explaining me this. Uh, and I call this fake helpful, and if you are able to put it as an imperative, it will save you a lot of time as we go. But you can clearly see we've, we've made a lot of tweaks over time.

    4. AG

      So for people who don't know, we showed Claude MD and Sol MD.

    5. MI

      Yes.

    6. AG

      What's the difference between those files? What should be in what? Sol MD I believe is a part of c- OpenClaude?

    7. MI

      Yeah, it's a part of OpenClaude. It gives, uh, e- effectively, uh, like, it's, um, a specific context that is being loaded into every single prompt. Um, and Claude MD has just the highest priority of them all, and Sol MD is the second in order of priority for OpenClaude.

    8. AG

      So your Claude MD, it seemed like you kept that under 100 lines, which I think is the advice that Boris Journey, creator of Claude Code, gave.

    9. MI

      Yep.

    10. AG

      The Sol MD, though, is, like, 800 lines, so that can be longer?

    11. MI

      Yeah, it's, it's longer, and, and you might argue that we are overloading the context, and some of those lines might even be contradictory. But what I've noticed was that that's the most robust, uh, way, uh, that gives me the most accurate output with the best recall. And we t- we test every single imperative. We actually test on, on the basic queries to see, uh, across the most important topics that we raise in the organization, uh, like h- how, how good it is or how bad it is.

    12. AG

      All right. What else do people need to know to set this up?

    13. MI

      So another angle is, um, you need to, uh, have tools, right? And it has a lot of tools. Um, so since it's interconnected, uh, with, uh, all of your, uh, workspace, it is deeply integrated into Google, so it's a separate tool. It's, uh, integrated into Atlassian. It's a separate tool. It has automations. An automation like the one that we currently have, it just consumes all of the mobile reviews. Uh, it writes a digest like a typical supportability team would do, and it even tags the people that are responsible, and it, it can even raise red flags in case it's necessary.

    14. AG

      And these are all Python files. You just prompt these with natural language in the IDE to Claude, and Claude writes these? Or how does that work?

    15. MI

      I previously, I used IDE, uh, but now I'm, I'm doing this for showcase purposes mostly. Um, uh, I'm, I'm using Claude App-

    16. AG

      Mm

    17. MI

      ... uh, because it allows me to do this, uh, on the go right here. It, it, it is accessible from my, from my mo- mobile phone in case if an urgent thing appears and I want to change anything within the agent.

    18. AG

      So do you basically create a GitHub repo that Claude Code can access?

    19. MI

      Yes, exactly. So every single change is immediately committed into GitHub repo.

    20. AG

      Awesome. So you do everything via GitHub via Claude App. That's so powerful. So you can configure, people kind of have this mistaken thing that they need to use the OpenClaude gateway, but you can actually do everything through Claude, and it can manage it on top of Hermes and OpenClaude.

    21. MI

      Absolutely. So the only benefit IDE gives you is that you can clearly see the, uh, the structure of the project.

    22. AG

      Yes. And you went through two IDEs actually. So you showed us Antigravity and Cursor. Talk to us about when you should be using the Claude App versus Antigravity versus Cursor.

    23. MI

      M- My use case are the following. So everything that lives in the cloud and, uh, doesn't require, I would say, um, access to my computer, um, I'm using, uh, I'm interacting with it through Claude App. Uh, but, um- Any specific use cases that I have that require my computer use, and mostly these are around, uh, browser, browser use, uh, via different MCPs or Excel use if I want to load a file and get some immediate feedback, or data research if I don't really want to load it into agent and I wanna do something ad hoc, I use, uh, an IDE for this. And why do I have two IDEs? Simply because it allows me to have two autonomous separate instances, uh, of agents running at the same time.

    24. AG

      Okay. I think I understand the OpenClaus setup part of this. It's mainly through the SolMD plus OpenClaus what's giving you that gateway to the Saul Goodman agent in Slack.

    25. MI

      Yep.

  15. 37:4641:43

    Hermes auto-skills + evaluation: measuring recall improvements and building evals

    1. AG

      Explain the Hermes part of this. You had said it's related to the skills and the recall?

    2. MI

      So, um, the beauty of Hermes as a scaffolding is that it comes with a unique advantage over OpenClaus, uh, which is it automatically generates skills based on the tasks that you most frequently request, uh, the system to perform. And what, what we've noticed in testing was that, uh, the recall has improved significantly. Well, I mentioned th-30% with those automatically generated tasks, uh, skills, sorry, versus without them. So we, we have it running under the hood all the time. Like for instance, team hiring evaluation, uh, immigration, like case building, because we have like different people who are relocating locally, so I have a lot of questions around it. Then it, it made a decision that we should have a skill for this.

    3. AG

      Mm-hmm.

    4. MI

      Uh, this, the same team hiring evaluation, it also made a decision, why do, why do I continue doing repetitive tasks? Why not just create a skill and offload all of this onto an agent? So i-it even kind of takes this meta part of the understanding, uh, whether or not you need the skill and makes a decision for you. That's, that's an amazing part of it.

    5. AG

      Awesome. And how did it help with the recall? You mentioned it also helped there.

    6. MI

      Plus 31%. So what, what does it actually mean? Uh, so, um, we have five core topics. Those topics are, uh, market, uh, its business model, uh, its product, key growth levers, uh, selection, uh, price, buyer activation, trust. Uh, then we have, um, more tactical things like funnel and so on. And, uh, um, we have a set of different questions that, uh, we ask. Um, and we ask like, I think like 10 questions for each of those areas, and we compared the non-skill response versus the skill response. And, uh, we also looked at how accurate the response in each of those categories for each of those questions was. And what we've seen was that having those skills in place gi-gave us plus 31% more accuracy-

    7. AG

      Wow

    8. MI

      ... which was, which was a deal breaker.

    9. AG

      How does someone set up like the metrics to measure something like recall?

    10. MI

      Um, you can do it, um, basically yourself. I'm, I'm doing this as a CPO because I think outside of knowledge graph and the percentage of context digitization, so, so to say, uh, you should also be able to track more tactical metrics, uh, which is recall. And you can ask, um, uh, Claude or whichever system you're working with or your agent directly, uh, "Tell me which areas are the most frequently, uh, asked or which I freq- most frequently interact with, and come up with 10 questions per each of those areas, uh, which are the most frequently asked." And then let's do a comparison, basically like one, uh, control group versus the treatment group with the skills and see what's the difference there. Um, that's, that's pretty much it.

    11. AG

      To use sort of a derogatory word, but not in a derogatory way, you're vibing the metric. You're, in natural language, you're explaining exactly how you want it to work, and then it's creating the metric, and it can go off and measure it for you.

    12. MI

      That is correct. But I mean, for the first pass, you have to like at least visually look at the results.

    13. AG

      Mm-hmm.

    14. MI

      Uh, but, uh, once you have a methodology set up, then you can delegate the evals fully to the system.

    15. AG

      All right. Amazing. So now people know how to set this all up. Can you show us some more use cases? What are the best, most powerful things people should be using this for?

  16. 41:4357:18

    High-leverage use cases: board-deck critique, permissions, prototypes, design systems, and hiring automation

    1. MI

      Uh, one cool thing, um, is this is, uh, this is also around skills, but I think this would be highly practical for people, especially in C-level positions. So I have a board of directors, like, uh, that I'm, I'm in constant touch with, and they make investment decisions, right? Uh, how much money, uh, we as a company should receive at which point in time, uh, what would be the payback, and so on and so forth. Um, and there are multiple people in the board, and they have different perspectives. But the beauty is, um, uh, the agent was able to abstract their mental models into a set of principles. Uh, and, uh, we called it this, the board skill. So in case if you have a pitch deck and you s- you wanna defend a strategy for the next year, then the first thing I do is I run this pitch deck through the board skill, uh, to, um, to kind of poke holes and give me brutal feedback on what could go wrong with the defense of this strategy.

    2. AG

      Fascinating. So your Granola meeting transcript is automatically recording your board meetings. Hermes has automatically created a skill around the profile of those board members.

    3. MI

      Yeah.

    4. AG

      And so when you're creating a board deck in Claude, let's say you use Claude Design and you would put in your input, towards the end of that process, you're gonna hit it with a prompt like Use the board skill to see how the board would respond so that we can refine this. Is that right?

    5. MI

      That is exactly right.

    6. AG

      Fascinating. That brings up a interesting point for me. We, as CPOs, we may not want to let our PMs see [chuckles] what's going on in a board meeting. How do we make sure that information is locked down so that the right groups get access to the right information?

    7. MI

      Oh, that's a, that's a great question, so. And, uh, two things. So the first one, it all boils down to, uh, scope of ownership. Uh, and, uh, the agent has, uh, different access rights depending on, uh, the person reaching out, and that access rights will define, uh, the context which this, uh, person or employee is able to retrieve, and the tools which this a- uh, which, which this person would be able to work with. For instance, board skill is available only to me and to, uh, to the executive committee within the company. Uh, and the second thing is the privacy aspect, right? Uh, because, like, not all people really appreciate that, uh, uh, their personal meetings are transcribed or recorded. So unless you willingly provide this information to Granola, we will not train the model and we will not create, like, the skills on top of this, so, uh, every, every employee has a say in this, of course. Uh, I'm, I'm personally a part of the experiment, that's why I'm completely transparent and I don't care about my privacy at all. [chuckles]

    8. AG

      [laughs] Okay. So if you slip in that you, uh, you know, were taking care of a sick kid during the weekend, you're okay with that hitting a transcript, but if a particular employee doesn't want it, they can elect to take it out.

    9. MI

      Absolutely. Absolutely, yes.

    10. AG

      Got it.

    11. MI

      And fr- from the get-go, we don't really, um, uh, we don't really store the transcript, the personal transcripts of people's meetings, because I think that violates privacy. So unless you, uh, willingly provide us this information, we will not do this.

    12. AG

      Mm. So you can select, like, this is a product trio meeting, this obviously makes it in. This is a product review, but this is just a one-on-one between me and my designer, this is not-

    13. MI

      Yeah

    14. AG

      ... gonna make it in.

    15. MI

      Yeah. Exactly, exactly.

    16. AG

      Got it. So Boardex is one really cool use case. What else are you, we- where else should people be using this for?

    17. MI

      Let's see. Uh, it responded in a different language, but, uh, the thing is, uh, so we asked it, "I want a specific feature," and instead he created, um, a prototype.

    18. AG

      Whoa!

    19. MI

      Let's see. Yeah.

    20. AG

      That's quite agentic and autonomous. [chuckles]

    21. MI

      It's agentic and autonomous, um, and it, it sometimes gets finicky, uh, depending on which language we interact with, and we use different languages, like, uh, so that's why it might get frustrated at times and it requires an imperative. But, but here, I think it created, uh, a prototype which is, uh, very close to our actual design system.

    22. AG

      Okay.

    23. MI

      Um, which is, I'm sup- I mean, it's not great in terms of UX, but, uh, uh, as a, as a one-shot pass, I think that's, that's an okay, uh, okay thing.

    24. AG

      As a one shot? I mean, it's just so much stuff it's done.

    25. MI

      [chuckles] Yeah, pretty much so. And it's very compliant with the actual Nexus design system, which is an OLX design system at play.

    26. AG

      And do we know what underlying model it hit? Did it hit Fable, or does it, do you have intelligent model routing? How does that work?

    27. MI

      No, it, it, um... So the thing is, um, uh, there is, uh, uh, an agentic orchestrator which makes a decision on which model to use.

    28. AG

      Mm.

    29. MI

      So if it's, uh, if it's a complex, complex request, uh, or if it's, um, uh, a specific domain area which, uh, has high sensitivity or error blast radius, most likely it will assign Fable onto it. Uh, if it's, um, if it's just an execution type of work or a status report writing, I think it, it will assign Opus 4.8 most likely. Uh, for low-level work that doesn't require super accuracy, it will assign Sonnet. So, uh, token optimization, uh, becomes, becomes, uh, important here.

    30. AG

      So mostly it builds on the Anthropic stack. You're not throwing in a Codex or a GLM yet.

  17. 57:181:06:07

    The future of PM and org design: smaller teams, blurred roles, and hiring AI-native PMs

    1. AG

      What's on everybody's mind right now, just in the zeitgeist, [chuckles] is what's gonna happen to my job? He just showed me how, you know, you might be able to do the work of two PMs with one PM. So what's gonna happen to the future? Is product management going to exist in the future? What's your take?

    2. MI

      Um, I have good news for you. I think, uh, not only product management is going to exist, I think it's gonna thrive. Uh, because I think, like, product managers in, product management in large organizations became, like, burdened with layers of hierarchy, with layers of corporate processes. And in, in many instances, product management is about compliance, performance, status reporting. So it's suboptimal product management, to be honest. And I think what AI really helps you to do, it doesn't remove away the decision-making from the process. Uh, it, it gives you the juiciest bit. It removes a lot of the operational overhead from your work. And, uh, what product managers are, um, now more focused on is the value discovery. Because, like, what the AI doesn't really know is, uh, what your customers want. Because the AI doesn't really understand the, um, um, the customer journey well. Well, it, it has the funnel metrics obviously, but it doesn't, uh, it cannot really go offline and talk to a customer. You have to do this. Um, and that's the funnest part in the job because you, you don't do all of this, uh, you know, um, theater anymore. You're only focused on delivering the real value. And I don't think it's gonna g- go away anytime sooner. What, uh, what actually will happen is that the teams will be smaller, leaner, and faster.

    3. AG

      So do you anticipate that the ratio of PM to engineer is gonna change from what it historically was?

    4. MI

      Yes, definitely. I think not only it's gonna change, it, the lines are gonna blur between what a product manager does and an engineering team or an engineering manager does. Uh, there's likelihood that, uh, they could overstep, uh, and, uh, be a support system. Because, uh, at the end of the day, if a product, uh, manager is an orchestrator of, um, of, let's say, the product quality, uh, and token budget, right? An engineer is an orchestrator of an execution quality and token budget, right? The same applies for the designer. Designer is responsible for the consistency of a design system, which means that quality and token budget as well. So, uh, the boundaries between those roles would be very blurry.

    5. AG

      So you're a CPO, you're thinking about maybe I need to create a new product team in a particular area. When I was a VP of product at Apollo, we would typically think about, okay, if we're gonna create a new front-end product area, we probably want five to seven engineers, we want a PM, we want a designer. When you're thinking about a new product area now, how do you think about staffing that?

    6. MI

      No, it's a, it's a great question. So I, I typically divide them into different groups. So the first group is, uh, high complexity and, uh, high error blast radius. I think for those specific domains, um, nothing has pretty much changed, and it's, uh, it's all around monetization. It's all around, uh, algorithm, algorithm like search, ML, uh, where every single, every s- uh, where a single tweak might have a dramatic impact on the conversion or on the outcome. Uh, here you would need to have a dedicated owner. Um, but for the remainder of the domains, like customer-facing domains, uh, what I'm currently seeing is that teams can be, like, easily scaled without adding people. You, you might have one product manager owning like three, four domains across many platforms at the same time.

    7. AG

      Okay. So the PMs you're hiring, they feel to me like- They're truly AI native PMs. Even if they're not building AI features, you're probably hiring people who are really AI forward. What's your process these days to hire PMs? How do you find PMs that are at this level of AI native that they can actually succeed in this environment?

    8. MI

      Great point. So, uh, I typically look at three things. I think the fundamentals haven't changed, uh, but I also added additional qualifiers into the interview. So the fundamentals meaning, uh, problem-solving and, um, systematic thinking. Like if you could have both, you will adapt in any situation. Uh, but additional qualifier that I've added is, um, a level of, um, ah, craft when it comes to AI, right? How deep a product manager is in the AI. And, and there, there is a very simple way to, um, test for this, is, um, I'm, I'm, I'm asking like, which of the use cases of your daily work you have automated using AI? And, um, I get a spectrum of responses. If a response is on the, on the sc- on, on a level of, well, I'm, I'm talking with ChatGPT using like, I don't know, terminal and using like web interface. And so to me, that's, um, kind of a low level of, uh, immersion or maturity when it comes to AI tool understanding. But on the other side, you might have, uh, a very advanced AI agent orchestrator, a product manager who has automated all of his professional work life, uh, right, and build it entire using AIs and so on. And then I can drill deeper, like asking like how do you evaluate this? How you do make decisions on tools, on the brain, on the memory, and scaffolding, and so on. But I gotta be honest with you, uh, it's, um, in the market where I operate, it's a rare skill set. So, um, product managers who are able to effectively answer those questions, they get way more points.

    9. AG

      You guys heard it from him, not from me, guys. He is a CPO at one of the greatest companies, I would say, to work on, and he's running it in an AI native way, and this is the skill set you need, the skill set we taught today. So if you haven't already, go play around with Hermes. Go play around with OpenClaw. And if you're a CPO, steal Mikhail's playbook. His team is operating at a level unlike many others. Mikhail, if they want to learn more about you, get in touch with you, find your content, where can they go?

    10. MI

      Uh, well, I do have a Substack newsletter. It's called Corporate Waters. Uh, you can reach out and, uh, get, uh, immersed in my thinking and, uh, the actual use cases. Um, outside of this, you can reach out to me directly over LinkedIn.

    11. AG

      I highly recommend Corporate Waters. If you guys didn't know, Mikhail and I did a really fun deep dive last year where he looked through all of the interviews he's done in his lengthy career, which by the way, before he was a GPM at Bolt, he was also an ICPM for over a decade at companies like Yandex. So he collected all that data, and we looked into what interviewers are looking for. So if you want to go see more information from Mikhail, you can find it in my newsletter or his newsletter. Thank you so much for sharing all this amazing sauce today.

    12. MI

      It was a pleasure, Aakash. Thank you for having me.

    13. AG

      The craziest thing is that Mikhail agreed to open source the information that he used to build all of this. So go check the link in the description for the GitHub repo on his profile. You can fork that, and you can begin building this company operating system for yourself. Until the next episode, we'll see you later. I hope you learned as much from today's episode as I did. If you can do one thing that's totally free that would help the show, it would be to check that you're following on Apple and Spotify podcasts. Check that you've left ratings and reviews on those platforms. Check that you're subscribed on YouTube. Leave a like and a comment on this video, and then share it with your friends. We're trying to make better and better podcasts. After two years, we think we've gotten something pretty good going. So let us know what we can do to make it even better, who else we should interview, and we will put on the best shows we possibly can. Finally, don't forget my offer for the bundle. You get an entire year of my paid newsletter, plus my favorite AI tools, Bolt.new, Airtable, Speechify, Descript, Magic Patterns, Linear, Dovetail, Arise, and Mobbin'. That's $27,000 worth of value for just $150. So check that out at bundle.aakashgee.com if it interests you, and I can't wait to share our next episode soon.

Episode duration: 1:06:16

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode 9W1m4S6VDU8

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.