EVERY SPOKEN WORD
20 min read · 4,402 words- 0:00 – 1:52
Why everyone’s quietly switching to Grok
- CVClaire Vo
I don't know if you know this, but the whole OpenAI versus Anthropic thing is old news. Everybody I know is secretly becoming a Groq boy. Yep, ever since Cursor was acquired by SpaceX for, I believe it was 60 billion American dollars, everybody's been really excited about the new products released from the Cursor team. And in fact, I'm hearing more and more positive things about the Groq models for coding and the Groq models in general. In today's episode of How I AI, we're gonna go through all the releases from the Cursor, xAI, SpaceX, whatever, the Elon cinematic universe of code builders and model builders. We're gonna go through all their products, and I'm gonna tell you what I think of GroqBot, Origin, the new GitHub competitor from Cursor, as well as the Groq model. Let's get to it. This episode is brought to you by Bolt.new, the AI app builder for people who have ideas and want to ship them. Most AI tools spit out code that looks great in a demo and falls apart the second you try to do anything real with it, or they lock you into their own platform with no real way out. Bolt is different. You describe what you want to build, a startup MVP, a landing page, an internal tool, a side project, and Bolt generates production-ready code in minutes. Connect Stripe or any other MCP, hook up your domain, and deploy it live. Founders are using Bolt to build businesses doing real revenue. Product managers are shipping prototypes their teams actually use. Designers and marketers are launching campaigns without waiting in line. Anyone can build, engineering can ship. Everyone wins. You just need an idea and a weekend. Check it out at bolt.new/howiai.
- 1:52 – 3:22
Grok Bot overview and setup
- CVClaire Vo
I'm gonna start with the most accessible product on our list today, GroqBot. In case you missed it, the Cursor/xAI team released GroqBot, which is a chat style agent in a desktop and mobile app that you can use to do knowledge work for you. It's very similar to an OpenClaw, but it's simpler, hosted, and easier to use. If you look at GroqBot just from a UI perspective, and I'm just pulling up the marketing site right now, we'll look at my GroqBot in a minute, it is definitely pulling from that iMessage terminal streamlined chat experience, and it's doing something I really love. Not everybody loves this, but I love a multi-agent experience. I want each of my agents to have a job. I want them to have a name. I don't want one agent to rule them all, and the team has really leaned into this ethos with GroqBot. Now, I've been testing GroqBot. I had a couple days early access, and then I've been testing it earnest over about the past week, and I think there are some pros to GroqBot, and I think there are some cons to GroqBot. But I do suspect that GroqBot is going to catch the attention of Cursor customers who want an all-purpose agent for their employees. So let's go into my GroqBot and show you a little bit about how it works. Now, as I said, this has this very, like, iMessage style experience. It's a left bubble and a right bubble, and you just chitty chat with your GroqBots.
- 3:22 – 4:30
My 5 Grok Bots
- CVClaire Vo
I have set up a couple of GroqBots in my own account. There is Prody McProd, my product manager GroqBot that's hooked up to ChatPRD. I have one of the out-of-the-box GroqBots commitment tracker that just keeps me honest with things I've told people that I will do for them, basically emails I need to reply to. I have a money maker bot that chases invoices and payments and sales deals. I have a case study buddy, 'cause I'm doing a bunch of case studies, and I have a GroqBot that monitors all my data for ChatPRD and tells me trends that I should pay attention to. So I think the magic of GroqBot, if I had to tell you one thing that is the magic of GroqBot, it is plugins. So I've always thought that Cursor was the best MCP client. I still to this day, when I wanna connect to different MCPs in my stack, would still default generally to Cursor. I thought the connector experience was really nice, and I thought the harness did a really nice job of talking to all these systems. They have brought that great plugin MCP experience into GroqBot, but there's one killer
- 4:30 – 6:07
The killer feature: multi-account connectors
- CVClaire Vo
feature, which is for any one of these connectors, whether it's Gmail or Slack or an MCP, you can connect multiple accounts to GroqBot. So if you have two Gmails, you can connect two Gmails. If you're in seven different Slacks, you can connect the different Slacks. This is something that Codex has not figured out, Claude has not figured out. No one seems to figure out that I have approximately a dozen email address and Slack accounts and different things that I need to log in, even if it's the same service. This experience of setting up multiple accounts per connector, huge, huge, huge benefit. You could, like, package it up and send it off. It would have product market fit with me. As you can see here, I have already four email addresses connected to Gmail. I have my ChatPRD one, um, and a couple other business ones that I'm working on. Having all four of these connected and being able to traverse all those accounts in GroqBot is a killer feature. Just thank you to the Cursor/xAI team for doing this. If you did this and this alone, I would be very happy, but they did not do that and that alone. GroqBot ships with tons and tons of out-of-the-box, uh, plugins. These are generally the ones that were available in Cursor as well. And so I've installed your basic productivity ones. I've installed GitHub, and I've installed the ChatPRD plugin, and really, I've been using GroqBot as a way to just traverse and use my data. So again, it's like a little bit of... A fancy MCP client,
- 6:07 – 6:41
Grok Bot’s virtual machine and how it actually works
- CVClaire Vo
but it does come with something special, which is every GroqBot comes with a computer. And so you can actually go into this virtual machine, and this is Prody McProd's computer. It's not very exciting, but it can use Chrome, it can use Terminal, and it has some files. And so again, this is like a little bit of a OpenClaw Lite, where it has a virtual machine. That virtual machine can use the web, and it has connectors and MCPs.
- 6:41 – 7:35
Experience overview
- CVClaire Vo
Now, what do I think of the experience using GroqBot? Well, I will say setting up a GroqBot, super simple. You can simply create a bot. It will start, uh, up that machine, it will start up a bot, and then you can just tell it what to do. Um, so this one just got stood up. "What's the first thing I want you on?" And I say, "I want you to be a family manager bot that keeps track of my personal email and personal calendar." Can't spell calendar. And again, it's gonna go ahead and say, "Great. I'm gonna find all your connectors. I'm gonna set it up, and I'm gonna learn myself." Really, really simple, streamlined onboarding experience, and it's gonna go ahead and give me pretty good responses back and do work on my behalf. It's pretty chat based, nothing fancy, um, but its, its genius is in its simplicity.
- 7:35 – 10:08
What I don’t love about Grok Bot
- CVClaire Vo
Now, what I don't like about GroqBot. Well, I don't like that it's so simple, and I don't like that I can't hack it. If you look at my OpenClaws, they are chaotic and technical and impossible to use and high maintenance, and I love them. I love them. Um, GroqBot, I have a little bit less control over. I manage them a little bit less tightly, and so I love them less. You know, like my children, the chal- the, the joy is in the challenge. And I will tell you, GroqBot has not challenged me because it works simply and it works out of, out of the box. The other thing I don't like about GroqBot is you don't get to pick what model you're working with. You really actually don't get to have that sort of like soul.md control over what GroqBot does, how it's configured, et cetera. And you know me, I just like to craft out, out, as if out of clay my agents. So I like that experience. I don't think everybody else does. So if you're looking for that like highly tunable, highly hackable, highly transparent setup for your agent, GroqBot is not that. Um, OpenClaw or Hermes or any of those is, is a little bit more like that. And then, of course, it doesn't run on my machine. I don't have control over it. It's a third-party system. Some good, some bad. That being said, it works, and the connector and kind of like UX affordances are quite nice. I would say just related to that, the last thing that I don't love about GroqBot is whatever model it's using, and maybe it's Groq, maybe it's... Sounds a lot like Claude Slop, is I just, no vibes. Just no vibes. The vibes are bad. It's like I don't wanna hang out with this person. Doesn't really make me cheerful. I have tuned my OpenClaws to be exactly who I wanna talk to when I wanna talk to them. This, I'm getting, like, too much of that, like, classic Slop, not this, not that Slop, terrible names for products like a scope knife. Um, just stuff that I would never say. And so I think the personality tuning and the voice tuning across all these harnesses are really important, but they're even more important if you are gonna do this multi-agent strategy where you're gonna have people name their agents and have them do work for you. A couple use
- 10:08 – 12:20
Grok Bot use cases and my honest verdict
- CVClaire Vo
cases that I think you could use GroqBot for, again, I've made-- I like to make my agents have roles. So I've made a PM one, I made a data analyst one, I made a deal desk one, I made a finance one. I'm making a family one. I think just thinking that way, what kind of teammate do I want, and then giving that teammate the tools it needs and giving it the instructions it needs is a great way to design your agents. Now, I would demo more, but I actually think this is all about workflow, so the product itself is super simple. I think the use cases are what make it interesting, and then the connector experience being top tier I think is gonna help the Cursor, its xAI, SpaceX team get their claws in the enterprise use case. So I'm really curious. Drop in the comments, tell me what your use cases are. Um, I will continue to share ones, and I will let you know if I decide I'm gonna migrate any of my OpenClaws over to GroqBot. If you want me to move over, SpaceX team, just make GroqBot way harder to manage, and then apparently I will fall in love with it and never use anything else. I'm gonna keep an eye on this, but TLDR, super simple, great connectors. Use your agents as employees, and if you like a bad time, stay with, stay with OpenClaw. This episode is brought to you by Jira by Atlassian. The Teamwork Graph in Jira delivers 44% more accurate agent results with 48% less token usage. That's a huge difference when working with AI coding agents like Claude, Cursor, Codex, or Copilot. The hardest part of shipping with AI isn't the code, it's the context. What's the right ticket? What did the spec say? What got decided in Slack? The Teamwork Graph pulls all of that from across your entire stack, from Jira and Confluence to GitHub, and feeds it directly to your agents before they write a single line. You assign the work, the agent gets everything it needs, and a PR surfaces when it's ready. No digging, no context switching, no broken flow. Same team, smarter agents. Try 'em free at jira.dev. That's J-I-R-A.D-E-V.
- 12:20 – 13:47
Cursor Origin: the agent-native GitHub replacement
- CVClaire Vo
Okay, the next release from the Cursor team is Origin, the GitHub replacement announced at Cursor's conference this year and finally released in early access today, which will be a couple days ago when we get this episode live. So Origin is in early access beta, and they've called it kind of codebase inside the app. So if you're in the Cursor web app or you're in the desktop app, you're really talking about codebase, which is, of course, the right primitive for something that is code hosting. Now, the whole value proposition that Cursor has put in front of us in terms of why they should build a GitHub replacement is they are basically building a n- agent-native GitHub replacement, which means that it's got all the Git primitives. It's got code, it's got diffs, it's got pull requests. It's got all that same Git primitive stuff, but it's built in a UI way, it's built in a UX way for agents to collaborate with, and in particular, of course, the Cursor cloud agents and the Cursor desktop app and, and, and the Cursor CLI. So it's, it's Git. We're all stuck with Git. We've agreed on Git. Not how we agreed on GitHub. Hmm, seems like the tides are turning. But Cursor's bet is that we are gonna want a agent-native GitHub, and that GitHub's not gonna get
- 13:47 – 14:59
What Origin actually looks like in practice
- CVClaire Vo
there fast enough. So what does that look like tactically? Well, Origin has repos. You can, of course, import GitHub repos. I'll tell you a little bit about my experience there. It has pull requests, which you can see. Um, agents can reply to comments. They can be assigned as reviewers, all those sorts of things. And there's a, a small set of extensions, um, CI/CD extensions, for example, like build preview branches in Vercel for those who want it. Okay, so what does that actually look like? Well, you go into the Cursor web app, and you click codebase, and then you basically import your GitHub repos. Very smart. Now, was not smart to release this the day that GitHub had a major outage, or maybe it was genius to release this the day that GitHub had a major outage. So I had some trouble with this GitHub import, but all you do is click sync from GitHub. You authorize your GitHub account, pick which repositories you wanna sync in, and then they are here. As you can see, you have kind of the same primitives. You have code, so this is your code view. You have pull requests, so you have pull requests, and then you have some settings.
- 14:59 – 17:42
Why I’m not switching from GitHub yet
- CVClaire Vo
What I would say is f- with the GitHub integration, it's really... it seems like it's just a wrapper on the GitHub API. So yes, in theory, this has been kind of redesigned. You can see here, like, there's some nice things about how Bugbot feedback has been shown or how Cursor feedback has been shown. It has some, like, smart little suggestions at the top that are, are nice. But at the end of the day, this is, this is just like GitHub with maybe less features, li- maybe less integrated. Much more integrated with Cursor, so I appreciate that, and I can see the vision. I can see the vision, I promise. But right now, if I'm using a GitHub-hosted repo and I'm not ready to go, like, all in on the Cursor ecosystem, there's just not enough here in this early launch to get me to pull over. And so, yes, I can, like, @Cursor here and ask it to fix the Vercel preview branch issue please. Why is it failing? 'Cause this is failing. I can do that. Oh, we got an error, so maybe I can't do that. Um, GitHub's having an outage right now, so maybe this is a GitHub issue, maybe this Cursor issue. But again, like, today's, today Origin, I'm seeing that they're setting the groundwork for getting you to get your repos into Cursor. They're setting the groundwork for Bugbot and Cursor to be first-class citizens in your repo. I'm sure there are some affordances on, like, the CLI side that make it easier for cloud agents to open PRs, and it's just gonna be just, like, a nice buttoned-up experience. But again, if you have invested deeply in GitHub, which everybody has, and you have your automations and you have your actions and you have your code owners and you have all this stuff, we're gonna have to see more from this experience, more from why AI native, why agent native matters, for us all to move over to this new product. Now, literally came out today. They're not even calling it beta. They're calling it early beta. So I have faith. I'm excited to see what this looks like, but again, just the current GitHub-integrated sync, it's, like, a little slower. I get the redesign, but it's not giving me life, and I'm just too embedded in the GitHub ecosystem. I need to see something wow from this in order for me to click
- 17:42 – 18:52
What would get me to move over
- CVClaire Vo
unlock. You know, and, uh, and I've spent an hour with it, two hours with it, not that much. I've messed around with opening PRs inside Cursor. Again, like, it's redesigned. It's synced to GitHub though, so it kind of feels like I'm being slow migrated over to Cursor and that they just want me to get, like... that they just wanna get their claws in me and then really wow me with this experience. But out of the box, I'm not yet feeling the vision, although I'm excited to see how this plays out. It's worth experimenting with. I think it's worth keeping your eye on. It's interesting. No one is happy with GitHub right now. Both stability is, like, terrible, and they really haven't yet given us that, like, AI native, completely reimagined Git experience that we're hoping for. That being said, it is super early on Cursor Origin. I think we are mostly in the very, very, very early stages of a very, very, very long migration, and it will be interesting to see if Cursor becomes the new centralized source of truth for code, or if it continues to just operate in the writing code agent's space.
- 18:52 – 20:41
Grok 4.6 and the How I AI Vibe bench
- CVClaire Vo
All right, so those are our two products. I would say I like GroqBot. I have not yet seen the light on Origin. The last thing we're gonna talk about is Groq 4.6. This is the model that has everybody I know texting me saying, "I think I kind of like the Groq models." Now, I think we're all trying this because Cursor has defaulted, I've noticed, some of their models to Groq. Um, not sure if I love it, but that means we give it a try. And, you know, we're gonna put benchmarks aside. Everybody picks the benchmarks that they like. I wanna talk about the How I AI Vibe Review. For anybody who is new, whenever a new model comes out, I do a couple things. I test PRDs, I test prototypes, I test design, wireframes, technical changes, and if I like talking to the model. I run all these evals blind, so I run them against models, and then I actually go through and grade all of these designs live myself. I've made some changes to the How I AI bench, in particular this design one, where I let the model decide how it wants to redesign the page, and I've made a eval for a very complicated claims adjudication wireframe, um, that has a lot of complex interactions. So those are two new things that we added to the How I AI Vibe bench. And then we ran Groq 4.6 through these models. I graded them. AI judges them, I judge them, and then we put them together in a presentation that I show with you before I even know what the scores are. So let's get to the How I AI Claire Weighted Index of Groq 4.6 compared to other models.
- 20:41 – 23:03
Claire Index results: where Grok 4.6 ranked
- CVClaire Vo
All right, so here are the results. Now, remember, this is my benchmark, so I get to decide how we grade things. And so it is 70% my taste, 30% our LLM as a judge taste. And interestingly enough, Groq 4.6 is right up there with my, my favorite, 5.6 Sol, beating out both Sonnet 5 and Opus 5 on the Claire Index. I am not surprised by this. I do love GPT-5.6 Sol. It's my default. But I am surprised to see 4.6 up here. I didn't hate it. Now, if you look at the overall recommendation by task, you'll see that 5.6 is my favorite direct writer on PRDs. I just like the clean way of writing. It's very matter-of-the-fact. It's very comprehensive, a little bit technical, but not overly complex. I also like a GPT-5.6 prototype, and so, um, I rated those quite highly. Opus 5 got graded by the LLM as a judge as the best implementer for a technical, like, bug triage problem. And then absolutely zero surprise, I continue to like Sonnet 5 for an OpenClaw agent chit-chat experience. This always wins. I know Sonnet 5 and for very pithy AI back-and-forth OpenClaw model, I just have not found anything better. Now, if we look at my opinions about prototypes, it's really interesting to see that when following art direction, I like 5.6, but when given broad decisions to make its own design decisions, I like Groq 4.6. Now, I suspect that this is because I can just spot GPT and Claude slop a mile away. A mile away. GPT-5.6 loves forest green. Claude loves a brown, tan, orange combo. So I think I was just enjoying this breath of fresh air that was Groq 4.6 that was not those two things. But in terms of executing complex UIs, I still really like 5.6 Sol, and when following art direction, I still like 5.6
- 23:03 – 25:00
Design evals: where Grok surprised me
- CVClaire Vo
Sol. Okay, so let's see what it actually shipped. So first, uh, question I was trying to answer in this benchmark is can it follow explicit design d- instruction? And GPT-5.6 Sol and Groq both followed my instructions well and did a relatively good non-slop job of following instructions. GPT-5.6 just continues to crush at this dock routing app that I have it make, where it has to give a very complex but easy to parse at a glance UI for managing, I guess, like, container ships. Crushes it, always does better than the c- competition. But I think that Groq did a nice job with my technical incident triage app as well as this editorial page. Now, when it's not given any design, we saw that every model did pretty good. We saw a couple good designs out of 5.6. We saw some clean designs out of Claude Sonnet 5, and my favorite one was actually this design from 4.6, which is Claude slop adjacent. We see an orange, we see a brown. But was actually quite cute and a good, um, interactive kind of ordering system for a coffee shop. So I was really pleased with that one. Dense information architecture, I, of course, loved 5.6. I think it's the best at complex UI and designing things that are not overwhelming, but are technically complete. And then in terms of just, like, taking a generic design and making it little better, I saw some good results from Opus and Sonnet. But when you aggregated all this up, all the ones I scored, good and bad, the two that bubbled up to the top were, again, 5.6 Sol, as well as 4.6.
- 25:00 – 27:13
My conclusion and how I’m splitting my time now
- CVClaire Vo
Now, what's really funny about this, again, this is 70% my taste, 30% the LLM as a judge. If you take my taste out of it, the LLM hates Groq, and it loves Claude Opus, it loves Sonnet, and it does not like GPT-5.6 Sol. Again, this is not weighted by an Anthropic model. I actually used GPT 5.5 to grade all these things, because it's actually the harshest judge. Um, so while the models still like Claude, I, Claire, like GPT-5.6, and I like Groq 4- Groq 4.6 for design. Now, this isn't my experience talking to the model, although you all know I love to do that. Um, this is pure outputs. So all I will say as my conclusion here is the Groq model is a competitor. None of it was terrible. I liked a lot of it compared to the models that you would use out of the box. And combined with Cursor's harness, its investment in new ways of thinking about things like Git, and fun products like GroqBot, I would say it's time for all of us to maybe become some level of a Groq boy. I am so excited about all this new stuff. It's really fun to look at. Again, true transparency, I'm still spending a lot of my time in Codex, I'm spending a lot of time with the 5.6 models. But given this experience, I might think about spinning up Cursor a little bit more for coding, and see if I can continue to like and get value out of Groq, Groq, Groq. Thanks for joining this mini episode of How I AI. I wanna hear from you whether you love GroqBot, if you think Cursor's gonna replace GitHub, and whether or not you've switched to Groq for any of your coding tasks. See you soon. [upbeat music] Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.
Episode duration: 27:14
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode 8ONFvAtboZ4
