Welcome to Product Cocktail, where the takes are as polarizing as a shot of Fernet—but the insights come together like a perfectly crafted daiquiri.
The Shake

And just like that... the summer is waning. The days are getting shorter, Halloween displays are creeping into CVSs, and cases of Shipyard Pumpkinhead are appearing at local liquor stores.
Back in May, I promised to keep you updated on my AI PM stack a few times a year. Looks like it's time for an end-of-summer AI stack round-up. (Beat drop.)
When I last wrote about my stack, I bemoaned how daunting it was to put it to paper (er, bits? tokens?). It was the best of times, it was the worst of times. There were a gaggle of unfamiliar-to-me vibe coding tools in the mix, Anthropic's product roadmap felt like a bullet train, and everyone (yes, including me) was building their own agents to... I guess solve world hunger or maybe just manage their family's sports pick-up/drop-off calendar.
I feel significantly more comfortable writing this version. Perhaps I'm jaded, perhaps there is genuine plateauing and consolidation, or perhaps this little experiment has deepened my AI awareness. You won't find breathless prose here glazing up every tool on the list (with affiliate links to discounted trials), you're getting my honest opinion, real use cases, and real failures—most of this issue is subscriptions I cancelled.
Let's dive in.

Four of seven categories changed in sixteen weeks. The three that didn't are the ones I'd bet on. (Source: Product Cocktail)
Daily Driver — Keep: Claude; Downgrade: Gemini
Claude Opus 5. Anthropic's productivity stack is the modern Microsoft Office (i.e., before they fell on their face trying to build collaboration). I've almost completely switched to using the desktop app.
It's much easier to open up a new Claude Code (or Cowork) session, pointed at a local directory with a pre-loaded CLAUDE.md instructions file, and pick things up where I left off vs. reloading context in a new Claude.ai web chat. I've also gone deeper into the feature stack, using /remote-control extensively to keep Claude Code sessions going on my phone. Particularly helpful to check in and respond to open questions if I kick off a long agent coding run before a workout or heading out of the house.
Conversely, Gemini—which I glazed up in my last tool stack post ("Gemini 3 felt like a big leap forward, on par or exceeding ChatGPT in many ways.")—has gotten worse through stagnation. At this point, I'm basically using it as a glorified Google Search proxy where I ask questions rarely requiring serious horsepower about cooking, translations, cocktail making fare, etc. Google hasn't shipped a flagship model in 189 days. Gemini 3.5 Pro was announced at I/O in mid-May for a June ship and has missed three targets — June, mid-July, early August. Compare that to OpenAI (GPT-5.6 Sol - 49 days ago), and Anthropic (Claude Opus 5 - 34 days ago) and it starts to look pretty grim.
I've started to use ChatGPT again occasionally—when I hit my usage limits on Claude. I'm genuinely impressed by the image generation (just don't use it to make posters 🤮).
Market Research — Keep: Claude; Downgrade: Perplexity; Add: NotebookLM (ish)
In May this section described a process, not a product: the same research prompt run overnight on Perplexity Deep Research and Claude, two cited markdown files in the morning, then a CLI agent pointed at both to reconcile them. The second model was insurance—mostly against confidently wrong numbers.
I don't run it anymore. Claude's citations got good enough that the second opinion stopped earning its keep.
Perplexity is the reason it's no longer a two-horse race. They've broadened the product to the point that I'm likelier to hit a nit—the maps experience, an aggressive push toward "Computer," no intuition about whether I want an answer or a webpage—than a research answer. I still use Comet interchangeably with Chrome. But when I actually need research, I reach for Claude.

"What do you want to know?" — followed immediately by four things Perplexity wants me to do instead. (Source: Product Cocktail)
I recommend NotebookLM (er, "Gemini Notebook") constantly, but rarely have opportunity to use it. When I have used it recently (technical interview prep, "reading" the Pope's encyclical), it has been critical in synthesizing and interrogating a specific set of documents with less fear of hallucination.
AI Meeting Notes — Keep: Granola
I'm still using Granola. My recent killer app has been mock interview analysis. It's a great way to get nuanced feedback and track development themes across interviews.
I've even played around with the Granola MCP as one of my first Model Context Protocol experiments in Claude Code Skills. There's a lot you can potentially do there with automating action items and follow-ups from meetings, or capturing product pain point themes from customer interviews. Pro-tip: use cheaper model sub-agents to read in and summarize transcripts, otherwise you're going to blow up your context window.
Rapid Prototyping — Add: Claude Design, v0; Kill: Lovable, Replit
This is an area where my expertise and opinion has dramatically shifted. As loyal readers are aware, I ran an extensive AI Design tool bake-off over the summer, comparing four tools (v0, Claude Design, Lovable, and Google Stitch) on three scenarios (change a screen, add a flow, design an app 0 to 1).

Three scenarios, four tools, one rubric. Lovable's 16.5 on scenario 3 is why the honest verdict is "mixed," not "bad" — and why Stitch's 7 and 8 are evidence of its ineptitude. (Source: Product Cocktail)
Coming out of that analysis, in the last few months, I've all-but-abandoned Lovable and Replit. Neither are bad products, and I'd happily use them again if given a corporate license. Lovable gave me mixed performance in the competition, sending me to greener prototyping pastures with Claude Design and v0.
Anecdotally, Replit is a bit token hungry (I spent about $38 to make the Linktree prototype — $18 above my $20/mo in credits) and a touch overbuilt for rapid prototyping.
I mentioned Claude Design's release in my last AI stack newsletter and it lives up to the hype. My interviewee Luke Moderwell put it succinctly: "Claude Design is head and shoulders above the rest." The same stack that can talk to your production codebase is now design-pilled, just watch out for the "house style" that can creep in. (You're not edgy for using Fraunces as a header font, you're just unoriginal.)
v0 took two of the three scenarios outright and held its own in the third. It's not cheap, but it can one-shot-prompt a very coherent design prototype, even just based on a screenshot of the existing product surface. I'm using the scenario 3 output as a starting point for my cocktail logging app, Proof.
AI Coding — Keep: Claude Code
Claude Code remains my primary coding agent, although I'm using it through the Claude desktop app now instead of Ghostty (terminal emulator). The integrated browser, iOS emulator, file directory, and artifacts list make it a far better experience than before.

The session list is the actual point: two Proof builds, and the rest is interview prep, newsletter scaffolds and pipeline updates. The phone on the right is Proof running in the simulator, inside the same window. (Source: Product Cocktail)
This is what I'm using to ship Proof.
Claude Code is also where I make model decisions, which is less fun than picking a daily driver and matters more. When I ran the extraction eval for Proof, Haiku picked the wrong menu entry 21% of the time. Opus—nine times the price per call—picked wrong 23% of the time. Paying more did not buy a better read of the menu. The thing that actually fixed it was a red box drawn by hand, which took both models to zero.
With these UI upgrades, it's a lot easier to use for broader, non-coding tasks (see below).
AI Chief of Staff — Add: Claude Code; Kill: Saul/OpenClaw
I'm a little sad to report that OpenClaw didn't really live up to the promise. On some level, this was a reality of the hype cycle moment it had and my inexperience defining agentic use cases.
Saul was extremely useful when I was getting my newsletter off the ground, bouncing ideas back and forth (especially while away from keyboard), drafting newsletter scaffolds.
I ran into tension with OpenClaw around two major things: token spend and time investment.
On cost: I didn't want to wire Saul in with frontier models, instead relying on Sonnet 4.6. This was good enough for most things. It still made me gravitate away from WhatsApp and more toward Claude Cowork/Claude Code where I had the power of Opus.
Time investment: In May, I identified five automation/enhancement use cases to deepen its value: Job Search Radar, Interview Prep Briefs, Newsletter Content Radar, Voice-to-Action via Wispr Flow, and a Weekly Accountability Pulse.
My scorecard is 1 of 5 on that set. I built a working Job Search Radar that scanned open jobs at a set of companies, applied a set of scoring heuristics, and sent me a daily message. It cost me about ~$1/day to run. The remaining ones, I didn't have time to build between taking care of a newborn, prepping for interviews, and writing this newsletter.
The primary newsletter thought-partner use case I mentioned above is easily replicated with a carefully crafted CLAUDE.md file, a local markdown directory, and /remote-control with Claude Code.
Back in May, I let Saul introduce himself in this newsletter. This is how the pitch held up.
No upsell to a $40/month plan.
Reader, I upsold myself. I throttled him to Sonnet 4.6 to keep the API bill sane, then spent most of my time in Cowork and Claude Code anyway—where I'm paying less for Opus than my monthly Claude API bills.
Most AI tools wait to be asked. I don't.
The one thing Saul ever did unprompted was the 8am job radar push. It wasn't compelling enough to keep running.
No context window amnesia.
This one held up. It just doesn't need Saul—it's a CLAUDE.md file and a folder of markdown.
One out of three, and the survivor defected.
AI Infrastructure — Keep: Obsidian, GitHub
The data infrastructure I built to feed my agent survives. It's just being used a bit differently now. Instead of Saul reading a private GitHub-synced version, I point Claude Code at a particular folder and away we go.
The Recipe

AI Models: Fanboying << Free Agency
Too much: Treating your AI model like a team allegiance. Dear Fable fans, Amodei isn't Tom Brady. Throwing the most expensive model at the task just because your favorite AI lab made it isn't good AI product design or intelligent workflow management.
Not enough: Testing on your own workload. Free agency between models is what deepens your AI product sense.
What's the fix?: Separate the two layers. Be loyal to a harness — that's where your context, your instruction files and your muscle memory live, and rotating through a new one every week is its own tax. Be a free agent on the model, where switching costs a config line. The one that does the job is the right one, and it isn't always the expensive one.
The Garnish

Two predictions. One winner, one loser.
Last stack issue, I showed cautious optimism for Stitch and bullishness on Claude Design. Let's see how those stack up.
Stitch:
Stitch. Google Labs' entree into "Vibe Design" (think: Figma but AI) launched a month ago and it has some rough edges, but it shows promise
In fact, it sucks and doesn't show promise of improving. It placed last in two of three AI design test scenarios. The one time a Stitch variant climbed off the floor, it was the pricier model—and it still only beat Lovable. It produces the fidelity of a wireframe made in PowerPoint when "clickable prototype" is becoming the norm.
Claude Design:
On the heels of Google releasing interesting-but-undercooked Stitch in mid-March, the race to replace Figma heated up as Anthropic dropped Claude Design. Now that the hype has died down, designers laud the speed but criticize “production-readiness.”
Right now, Figma and Stitch / Claude Design are complementary tools, but I don’t expect Anthropic’s product team to sit around on their asses with this one.
Banger product, with Claude Code integration. It earned a place in my stack.
Source: Product Cocktail
Product Cocktail
Tip Your Bartender
Reply with your own stack—what you actually use, not what’s on your LinkedIn skills section. If I'm impressed by your cocktail, I’ll feature it in a future issue.
Icons made by Icongeek26 from www.flaticon.com.
