Staging environment

Vibe Check for PMs: Opus, Fable, GPT-5.6, Kimi K3

Hosted by Aman Khan and Eric Xiao

637 students

What you'll learn

What to test after you upgrade to a new model

Which model is actually best for PM tasks

The specific quirks of Opus, Fable 5, GPT-5.6 Sol and Kimi K

How to quickly create evals to benchmark new models

Why this topic matters

Four new frontier models shipped last month with stellar benchmarks, but how do they stand up to real-world usage? Which models work best as a thinking partner vs. a task executor? We test all four against real PM tasks live, show you how to prompt each one differently, and build an eval you can use to benchmark performance for your own tasks.

You'll learn from

Aman Khan

AI PM at Arize, ex Spotify, Apple, Cruise

Aman has worked as a product leader at Arize AI, Spotify, Cruise, Zipline, and Apple. Currently, Aman is lead PM at Google on the Agent Platform. At Google, Aman helps teams launch and improve their AI systems. He recently led a popularΒ deeplearning.aiΒ course on Evaluating AI Agents, and has been featured by Lenny's Newsletter to cover AI Product Management a number of times.

Eric Xiao

Founder @ Bloom, prev. AI PM at Meta, Arize

Eric is a founder and full stack builder. He is currently building Bloom, anΒ investing assistantΒ (100k+ downloads). He previously led product at anΒ AI evals company, holds a patent for anΒ AI shopping assistant, and was an executive for aΒ series C AI startup. Eric uses AI to prototype, market, and ship new products from scratch every day.

See all products from Aman

Go deeper with a course

Claude Code and Codex for Product Managers
Aman Khan and Eric Xiao
View syllabus