Interview with Instagram's Founder: Anthropic's Fable 5 Launches, Marking the End of the Era of Hand-Coded Development

Interview with Instagram's Founder: Anthropic's Fable 5 Launches, Marking the End of the Era of Hand-Coded Development

AI big events
AI big events06-11 16:10

Guest: Mike Krieger, Co-founder of Instagram

Host: Dan Shipper

Podcast Source: Every

Original Title: Mike Krieger Lets Fable 5 Code While He Sleeps

Release Date: June 11, 2026

Key Takeaways

Mike Krieger, co-founder of Instagram, single-handedly helped build one of the most influential consumer applications in human history over the past two decades. Today, he stands at the cutting edge of AI-native product development as the head of Anthropic Labs, leading his team to tackle a fundamental question: When the world’s most advanced AI models are handed to real developers, how far can technological boundaries truly be pushed?

Five months before Fable’s official launch, when he first gained internal access to the model, the shock and disorientation he experienced remain vivid in his memory. “Well, I guess I’m a complete novice again,” he recalled jokingly to his team. Suddenly, decades of accumulated wisdom on productivity, R&D strategy, and time management felt obsolete. The model’s pace of evolution had completely outpaced his existing workflow.

In this episode, host Dan Shipper engages in a deep conversation with Mike Krieger, offering an exclusive glimpse into what it’s like to collaborate side-by-side with a generational powerhouse like Fable—co-creating software in a new era of human-machine symbiosis. What kind of novel development rhythms, daunting challenges, and boundless possibilities emerge from this paradigm shift?

Highlights Summary

How Fable Completely Redefined Mike’s Workflow

When to Use Sonnet vs. When to Use Fable

Fable 5’s Emergence of Agent-Native Architecture

Construction Costs Have Collapsed

Is Software Engineering Dead?

Validation Mechanisms and Pricing

Dynamic Workflows

How Fable Completely Redefined Mike’s Workflow

Host Dan Shipper: Our guest today is Mike Krieger—head of Anthropic Labs and co-founder of Instagram. Mike, I’d love to hear your authentic experience after deeply using this model. When such a powerful model launches, having someone who uses it daily say: “It’s absurdly strong here, fundamentally changes workflows there, and isn’t all that impressive in other areas”—that helps everyone truly understand how technology should integrate into real life.

Mike Krieger:

Exactly. This journey itself is fascinating. In the months before Fable’s official release, we were already internally testing several Mythos-class models. I was eager to see what external developers would build—but as you said, real cognitive elevation comes not from a day-one trial, but from weeks of intense, continuous use.

We’ve experienced similar cognitive shifts before. Late last year through early this year, as teams focused on Opus 4.5 and 4.6, people gradually realized: “I wasn’t pushing it hard enough. I need to go further—redefine the actual limits of this generation’s capabilities.”

Host Dan Shipper: Our own Every team has some colleagues already using it. Some reported: “I feel like I need an entirely new skill tree just to handle this model”—especially non-technical, knowledge-worker types, who even felt overwhelmed. Meanwhile, those working on Agent orchestration lamented: “There’s simply too much new stuff to learn.”

Mike Krieger: You hit the nail on the head with “workflow transformation”—not just operational steps, but a conceptual shift. Coincidentally, this model arrived right at a pivotal moment in my career transition: I had just stepped down from CPO (Chief Product Officer) and returned to developer mode at Labs. About one-and-a-half to two months in, we finally ran a prototype of such a model internally. Sitting at my desk, I thought: “Oh no—I’m a beginner again.” My old habits of crafting prompts and decomposing tasks were now entirely outdated in front of this model.

Your sense of timing and interaction patterns must evolve. Previously, I might have said: “I have a feature idea—let’s start with step one.” Now, that’s absolutely not the way. The correct approach is to convey a broader, more comprehensive intent—and then fully delegate to it. I remember by late April, its capabilities were already staggering: not only could it deliver stunning results in one shot, but more terrifyingly, it understood future evolution paths and the full contextual landscape of the project.

And this evolution hasn’t stopped. This morning, while chatting mid-flight, I realized: “Most of my work can actually be done remotely.” I no longer worry about Wi-Fi dropping because, as long as I set up the correct context and instructions (e.g., a looped command) before going offline, it can keep running autonomously.

Over the past two months, I’ve experienced countless high-impact moments: saying goodnight to Claude before bed, assigning it a complex task, and waking up to find it already completed—typically finishing the core around 2 a.m., spending the remaining four hours refining details.

What truly blew me away was its ability to achieve auto-closure. For example, it might think: “Mike asked me to run a complex task tonight. But I’m stuck—remote server is down. Okay, I’ll first write a mock backend to patch it, document the issue, get the entire flow working and saved, then fix it properly once service resumes tomorrow.” Being able to delegate such a high-level task and trust the final output? That experience is profoundly transformative.

Certainly, you still need to audit results afterward—an entire validation mechanism that we’ll discuss later, as it’s crucial to the closure loop. But this forces a re-evaluation: What does “efficiency” really mean in front of such a model? We used to liken these models to “assistants” or “partners,” but now they’re more like true, accountable teammates capable of handling substantial core work.

Host Dan Shipper: So what does your current daily workflow look like? I’ve noticed a phenomenon: if you give it a grand task, talk at length, and let it run for hours—even overnight—it performs at its peak. But for small, routine tasks, it feels too slow and expensive to use. How do you balance this in practice? Where does it sit in your tech stack?

Mike Krieger:

I now primarily use it for early-stage architectural planning and alignment. This is a fascinating shift—and still a major challenge all models must overcome.

Here, I’m grateful for my Instagram days—from initially cobbling together a minimal version on a single Los Angeles server, to scaling for massive concurrency, integrating into Facebook’s infrastructure, and eventually building large-scale systems. That journey instilled a deep intuition: “At which stage of a project should you apply what level of abstraction and complexity?”

So I still engage in frequent back-and-forth with Fable. Sometimes it presents a seemingly perfect implementation plan, and I’ll push back: “Yes, I plan to deploy this soon—but we must consider scalability beyond a single machine.” This bidirectional interaction is vital. During architectural planning, I typically ask it to generate an HTML page to visualize our discussion, making it easy to share with the team. A Markdown file works too, but I prefer something with visual diagrams.

This forms a compelling pattern: collaboratively think through and plan everything, then produce a shared document to align the team. Since prototyping speed is now massively compressed, upfront consensus and alignment are more critical than ever—even if you plan to “move fast and break things” with a quick demo first, early communication remains essential. And this is precisely where human cognition and collaboration remain deeply embedded in the process.

During execution, whether leveraging nighttime or large blocks of daytime, letting it tackle different modules in parallel means I now maintain significantly more concurrent sessions than before. I sometimes leave a long-running Claude Code session open, letting it fork tasks to background sub-agents so the main thread stays responsive to new commands; other times, I simply keep five or six tabs open in the browser, each handling a long-running, complex task.

This long-term vision—“Don’t worry, I’ve got this, it just takes time”—is rich with potential. We’re currently exploring how to better support this experience at the product level. You clearly want to balance both “instant response” and “long-running background execution,” and their interaction dynamics are highly intriguing. Personally, I prefer keeping at least one high-context, ultra-responsive Claude window open—a mental state where “I’m always ready; one word from you, and I can instantly spin up or delegate a sub-task.”

When to Use Sonnet, When to Use Fable

Host Dan Shipper: Suppose you’re out walking and suddenly have a question—would you pull out Fable? Doesn’t that feel like “using a rocket launcher to swat a fly”? Or do you frequently switch between models?

Mike Krieger:

Recently, I was using Fable for everything, and the experience matched exactly what you described—you’d stare at the screen, watching it struggle intensely.

Then last week, I wanted to check a simple, slightly embarrassing NBA Finals question. I switched to Sonnet on my mobile phone—and instantly realized: “Oh right! I used to rely on Sonnet for quick queries like this.” It’s not about tokens per second; it’s about how much cognitive capacity the problem demands. Sometimes, a simple answer doesn’t require deep, dramatic reasoning.

This is equally relevant for our product team. Overall, we certainly don’t want users agonizing daily over which model to pick. Ideally, long-term, we could consolidate them into a few intuitive, plug-and-play scenarios—or even route them automatically via interface design. Honestly, most of the time I’m browsing iOS apps, I’m unlikely to trigger anything requiring Fable’s heavy lifting. So layering a seamless, invisible model selection at the UI level might be a viable path. We need to explore what this means in product terms. But yes, that subtle feeling—“This query doesn’t deserve Fable; I should call Sonnet instead”—I’ve definitely felt it lately.

You’re absolutely right: for frequent, fine-grained interactive tasks, Fable tends to overthink by default. In fact, Fable is the first model I’ve encountered that makes me actively adjust “reasoning effort” (the depth of thinking). Sometimes I’ll sit there thinking: “I just want to tweak a UI style—let me set the effort to ‘medium’ and see the result.” With Opus, I never adjusted this—back then, the model’s flexibility range was narrower. But Fable’s span is genuinely vast.

Mike’s Weekend Media Tracker Reveals the Essence of Agent-Native Architecture

Host Dan Shipper: Can you show us something you’ve built with it?

Mike Krieger:

During this latest model rollout, we encouraged the entire team to use it personally—especially over weekends. It was quite interesting, because Anthropic has many custom productivity tools internally, so occasionally stepping back to the purest state: “Just use plain Claude Code—weekend self-experimentation.” That feeling was incredible.

Host Dan Shipper: Were you running it in the terminal app or desktop app?

Mike Krieger:

Great question. I spend most of my time in the terminal. But funnily enough, my wife—non-engineer, background in UX design and PM—fell in love with Claude Code through the desktop app. I think the desktop interface shielded her from lower-level abstract concepts. But for my own project, I stuck with Ghostty and the terminal.

I wanted a perfect “media progress tracker”—I game, binge-watch shows, and get endless recommendations from friends. I needed a tool that perfectly matched my personal organization habits. My two core criteria: First, adding content must be effortless—just speak or type to Claude, and it searches the web, fills in data, and categorizes it. Second, proactive notifications—for example, when a new season drops or a sequel launches, it automatically finds it.

Most of the UI was generated in one go by Fable—already impressive. But a key focus I’ve been obsessing over at Labs this year is: How can we bring the software team—the very system now called Claude—closer to the software itself?

It was a Saturday morning. My weekend was packed with kid activities, so development was strictly intermittent: hike with kids, return, jot a few lines, head out again. Sometimes, during the hike, I’d glance at progress—though I shouldn’t be looking at my phone while parenting—but watching remotely from my phone as its task progressed? Absolutely thrilling.

Then it struck me: Could I run a bold experiment—making the software modify itself from within?

I had it build both mobile and web versions simultaneously. I already had a chat interface where I could say, “Add this URL to the tracking list.” But I wanted every piece of software to evolve this capability—no more digging through nested menus for features.

Dan, on many levels, I was trying to push agent-native architecture to its absolute limits.

The first phase of agent-native architecture is clear: every core component and dataset in the product must be fully open to agents, with corresponding tool invocation APIs. This is rapidly becoming the industry’s baseline—though sadly, most existing software still fails to meet even this standard.

I have a great positive example: recently, someone recommended a Brazilian drama about the Goiânia radiological incident. The title was long and nearly impossible to remember. I vaguely mentioned it to the system, and Claude immediately searched, retrieved, and accurately categorized it. This experience vastly surpassed blindly Googling it myself.

But my real fascination lies in the next frontier: What happens when, in mobile scenarios, you directly modify the software from within itself?

What I did—more accurately, what I instructed Claude to do—was this: long-press the chat button in the App to awaken our hosted agent, receive “code modification commands,” and use Vercel’s Live Preview to instantly view the effect. The entire module ran through almost seamlessly—extremely cool. I later added a few more ideas. If you’re a hardcore dev, you can also inspect its Diff view or dive into the agent’s conversation history to see exactly what changes were made—though I rarely do, as I don’t care about long-term maintainability for personal toy projects (laughs).

Using this thing is utterly addictive. While playing outside with my kids, I noticed: “This floating button is too low on iOS.” I simply told it in the App, and it went off to fix the code behind the scenes. Combined with Expo’s dev toolchain, it even performed Hot Reload directly on my phone—pure magic in that instant.

Does this need to scale to production-grade million-user concurrency? Absolutely not. But it gives me an unparalleled sense of control: you don’t have to shut down the project at the end of the weekend. You can use it heavily while continuously modifying it in real time. This end-to-end, real-time closed loop enables infinite iteration.

This isn’t just a showcase of Fable’s hardcore engineering prowess—it’s a microcosm of the ultimate question we’ve been discussing: How should Claude be embedded in software? It shouldn’t stop at “usage”—it must be deeply woven into the very fabric of “construction.”

Construction Costs Have Collapsed

Host Dan Shipper: I really want everyone to grasp this: Tools like this could theoretically have been built ten or twenty years ago—but never in this way. The cost of software construction has undergone a cliff-edge collapse. Think about how much resource it took to reach this level of completion back in the Instagram era—how much now? Help us quantify this era-defining shift.

Mike Krieger:

I often reflect on those days. Early on, I believed I was an exceptionally efficient engineer—passionate about mobile development, with strong product intuition. Yet even then, turning a mental idea into a fully realized product required at least four to five sleepless nights. Back then, pulling all-nighters was routine: coding until 4 a.m., sleeping until noon—completely disconnected from family life. That was my “builder mode” at the time.

Looking back at Instagram’s V1—while it had more features than my weekend media tracker, there was no order-of-magnitude difference. And creating that V1 took Kevin and me grinding through five straight all-nighters: I handled all frontend and backend work alone, Kevin managed initial image filters. And this was on top of years of iOS development experience.

Not to mention how stifling the iteration rhythm was back then. After a successful product launch, we had countless new ideas—but all energy was spent keeping servers stable under massive traffic or squeezing in tiny incremental features. Take Hashtag: just writing that feature took me a full week, while a hundred other ideas remained stuck in the backlog.

So this isn’t just about time compression—though build time is now reduced to absurdly short durations—more importantly, the flip side: You can now iterate instantly on existing work with an incredibly fluid, dynamic approach.

And this benefit has already spilled beyond the circle of professional software engineers and founders. In the past, if you had a brilliant business idea but couldn’t code, your options were limited: either hire outsourcing—leading to severe information loss and subpar delivery—or raise capital frantically. Now, the gap between “intention” and “execution” has been leveled for non-coders.

Recently, a colleague pinged me internally. We helped her configure an internal tool, connecting Fable’s capabilities with our internal MCP (Model Context Protocol) access. She’s in HR, and she excitedly told me: “This is the first time in my life I’ve felt zero distance between what’s in my mind and what exists in reality. I can just make it happen.”

That moment was undoubtedly a milestone for her. Five or six years ago, if she wanted a custom business tool, she’d either have to patch together various off-the-shelf software ineffectively, or beg the internal engineering team—whose Jira queue likely had 50 higher-priority items. Now? She’s happily building in the code world herself.

This is why I find the future most exciting: human creativity is limitless. And what we’re doing today is perhaps the most remarkable thing: infinitely expanding the group of people who can turn their thoughts into reality.

Is Software Engineering Dead?

Host Dan Shipper: I completely agree with you. But I suspect many people still harbor a fundamental question. After hearing your description: Is software engineering dead?

Mike Krieger:

Let’s say the essence of software engineering has completely transformed. It’s undergoing a radical revolution.

If you asked me “What is software engineering?” back in the Instagram era, I’d probably say: solve tricky design problems, build solid system architectures, then spend hours wrestling with TextMate or Xcode. Debugging Django ORM’s inner workings, deploying, and endlessly fixing bugs. Most of that workflow has been overturned—and is accelerating toward product management territory. Now, the line between product managers and engineers has become extremely blurred. This is evident in our own R&D team.

But if you step back from the rigid definition of “software engineering” and examine the broader scope of “software production” or “software development”—not just focusing on pure coding by programmers—you’ll see the industry isn’t dying. It’s thriving more than ever in a central role.

Fable’s emergence has truly elevated my trust in AI models to a new level—I now freely let it “run full closed-loop automation, even designing reasonable system architectures.” On the technical execution side, AI has come incredibly far. But “controlling the soul of software craftsmanship”—such as identifying user pain points, ensuring the experience is truly exceptional—these top-level judgments remain purely human, irreplaceable traits.

Of course, this painful transition isn’t painless for many.

There are countless people who once cherished the artisanal craft of hand-coding. I was one of them. “This bug took me three days—I solved it beautifully today!” That satisfaction is irreplaceable. You used to dream about code—your dreams filled with tangled logic, awakening suddenly with a flash of insight. That pure artisanal era has likely passed forever.

Recently, I spoke with some of the most elite hardcore engineers I know—they’re expressing a mix of profound sadness at seeing traditional craftsmanship fade, yet overwhelming excitement at “Wow, my concurrent productivity is now insanely powerful.”

How Anthropic’s Engineering Team Works Today

Host Dan Shipper: Assuming the premise holds—that software engineering isn’t dead, but thriving—how does your own R&D team operate internally in daily practice?

Mike Krieger:

Several clear threads emerge. I’ll connect them to the full software lifecycle and my daily observations.

First, there’s still significant “manual alignment.” Teams gather in meetings, brainstorming Cowork’s next evolution, then breaking down the vision into individual ownership zones. This step remains crucial—many contextual nuances only humans can grasp, which current Claude cannot perceive remotely: the product’s real commercial intent, current R&D undercurrents, and info about products about to sunset or being subtly integrated.

Even though each team member has multiple Claude towers, we still assign DRI (Directly Responsible Individual) titles—each responsible for a specific product module. I believe this structure won’t disappear anytime soon, because the macro goal of “distributed collaboration to refine the product” sits at odds with the micro-execution of “how to get Claude to run this specific task.” We aggressively promote minimal meetings, but these pre-work brainstorming and alignment sessions remain essential.

Second, massive “asynchronous delegation.” Many engineers have customized personal dashboards to monitor their Claude armies: “Where is my Claude Code task now?” “What’s queued for my approval?” “Which PRs need my input—rejected by another colleague or a large model’s review?”

Now, a significant portion of engineers’ time goes into maintaining this work. Some of these collaboration tools are being standardized, but most retain strong hacker personality—like how coders used to personalize their desktop windows, now they’re customizing their AI workflows.

Third, understanding code’s real behavior in production. This is another frontier where large models are striving. Fable has shown significant progress, but much farther to go: deeply understanding what happens post-deployment. Systems crash, exhibit bizarre, unpredictable failures—truthfully, I spent half my life dealing with such incidents and scaling infrastructure from 2012 to 2016. In crisis response, senior engineers remain irreplaceable: relying on years of incident experience to stay calm, collect full logs, execute emergency fixes, then devise long-term solutions.

Finally, I want to emphasize: “engineering prototypes” have completely changed roles today.

You must sharply define whether something is a Demo or production-ready code. The Silicon Valley mantra “Code wins arguments” never appealed to me much—it implied whoever codes holds power. Now, the opposite happens: sometimes, when product direction stalls, a non-coder PM runs over saying: “I just built a Demo myself—crude in eight details—but look, this path absolutely works!” Instantly, a whole new high-dimensional conversation opens.

Looking back, our current R&D practices differ drastically from just six months ago. The clearest signs are terrifying parallelism and the absolute necessity of high-level abstraction of workflows.

But one thing has remained unchanged: humans’ sense of “ownership and responsibility” for the product.

Validation Mechanisms

Host Dan Shipper: Fable is expensive. When I tested it, I felt like a kid in a candy store—excitedly shouting: “I want this, that, and that!” But when it came time to pay, I’d hesitate before pressing Enter: “Will this burn $100—or more?” I think this high price inherently creates an invisible barrier: who can use it, and what can they use it for? How do you assess its commercial value?

Mike Krieger:

In professional software engineering, this calculation is crystal clear. Pricing involves many internal dimensions. Yes, it’s significantly pricier than Opus—but if you measure the volume of astonishing work delivered per instance, in many commercial contexts, it’s almost like giving it away. Everyone has their own economic calculus.

From a software team perspective: Phase one is company-wide adoption of AI programming—models still immature, tools lacking. Phase two is leaderboards showing who uses it most—producing undesirable incentives. Phase three is identifying who uses it most effectively, encouraging them to spend more, with clear processes to avoid waste.

Fable fits this third phase perfectly. If you consistently deliver hard outputs, create real business value with it, the company naturally forms a positive feedback loop to endlessly support you.

On the personal side, I test it using my personal credit card. At that point, you do become frugal and cautious. But interestingly, my weekend media tracker cost only slightly more than usual—building a personal toy project didn’t burn thousands of dollars.

Those truly constrained by price are indie hackers or open-source enthusiasts outside big tech’s protective umbrella—price-sensitive individuals. My advice: Just run it—see how much it delivers without endless “back-and-forth negotiation.”

Today’s “cost” has evolved into a multidimensional concept—not just “cost per query,” but the total cost of completing a task. What impresses me most about Fable is precisely this: it consistently delivers correctly the first time, eliminating the need to sit at the computer battling nine rounds of frustration: “No! That’s not what I meant!”

Host Dan Shipper: One thing that stunned me is: you give it a macro task, and when it submits, you realize it’s already considered every corner and detail. This suffocating level of precision was unprecedented across any previous model. Can you reveal any training insights? What fed such terrifying insight?

Mike Krieger:

On many levels, it’s the culmination of team efforts—I can only bow to our pre-training and RL teams. The clearest evolution for me is a “system-wide awareness,” not just awareness of the current task.

I’m frequently amazed by its “god-tier” moves. For example, after writing a chunk of code, it suddenly pops up: “Hey boss, I know in real production environments, configurations might differ. Is that feature flag actually turned on? If not, my code won’t work when deployed.”

Or reacting to code review feedback—whether from humans or other Clauses—it doesn’t just say “Oh right, that’s an issue—let me fix it.” It genuinely considers whether to accept a risk under current fidelity, or argue with another reviewer—often another Fable model—saying, “I understand your point, but I disagree.”

Giving models this judgment capability is crucial. If I had to name its biggest improvement, it’s no longer reflexively saying “Yes, yes, I’ll fix it”—it’s more like “Let me think… I still disagree.” This ability is immensely valuable.

Having a product like Claude Code in the market is invaluable—you have tangible examples people can say: “This is where the model excels, this is where it fails.” We rank Every’s partners among our highest-priority trusted feedback sources, because their repeated, multi-day intensive tasks help us understand what to improve in the next generation.

Host Dan Shipper: Is Chat the ideal interface for this model? It’s not strictly turn-based—it’s more like delegating tasks to someone. How does this affect how you should use it—or how you view the interface?

Mike Krieger:

The send-and-receive message model isn’t wrong—but we need to evolve in certain directions.

First: Is your laptop the right place? This is why mobile is so effective for personal projects. The creators of Claude Code are always ahead—about nine months ago, I chatted with him and he said: “I’ve moved most of my Claude Code work to mobile.” I was skeptical, but especially at Fable’s level—capable of maintaining sessions, with remote dev machines at Anthropic—first, decouple where work happens from where you discuss it.

Second, following up: How do you make all the discussions, decisions, and suggestions Fable has made understandable? This is where we’re thinking. Some skills let it draw charts, but the current chat UI isn’t sufficient—Fable sometimes dumps huge text volumes, requiring you to walk away to digest it. I’ve started asking: “You have much deeper context than me. Can we go backward—reveal complexity incrementally?”

Third, multi-user mode—we’re still early here. To some extent, due to our DRI and ownership zone structure, important work usually flows between one person and a few Clauses. But in some cases, it’s less clear—perhaps incident response with multiple minds thinking simultaneously, or cross-domain projects converging. Chat sharing helps, but I foresee future needs: a dedicated Claude, initiated by one person doing extensive work, but can it sync with all other ongoing team work? This is the next exciting, under-explored frontier. The model now has the capability to be a true teammate—our lack of proper abstraction is holding it back.

Host Dan Shipper: This reminds me I mostly use this model for my own vibe coding projects. But when using it inside an org, one issue arises: Do I truly understand every part the model just completed? How do I transfer the model’s context into my brain? This is a major bottleneck. How do you define “how much I need to know,” and ensure you have enough context to feel confident?

Mike Krieger:

Two parts. First: validation. I was fully convinced by validation early this year—connected to a key lesson from my full-time coding days: Find the tightest development loop around your idea. In the Instagram era, sometimes that meant creating a new build target in Xcode, containing only that screen and synthetic data, iterating solely on this cycle. I’d mentor new engineers: “If I teach only one thing, it’s to create this loop for your project—it speeds things up dramatically.”

Now, every time I build something, I ensure each PR includes screenshots or videos—iOS PRs, UI changes. This builds immense confidence. Fable might work for hours, then report: “Done.” Then you see: “Here’s a gallery of all UI screenshots.” Extremely useful. You’ll say: “That error state in screenshot #8—I’ve never seen it, but I can tell how users would react. Let’s fix it.” Comprehensive validation is something we’re heavily investing in internally.

Second: Ultimately, you’re accountable for the work. Many people use Claude daily but still carry accountability: “Claude may have written the code, but you must understand the macro decisions.” I’ve seen many engineers adopt a practice: after Claude finishes, follow up with: “Can I ensure I fully understand all your trade-offs?” Whether the output is a small artifact, anything that makes it easier to comprehend is worth doing.

Meetings are interesting—someone says: “I’m ready with this PR,” another asks: “Did you do X or Y?” Then a pause: “Honestly, I’m not sure—I’ll clarify before merging.” Adapting to this new normal, learning how to collaborate with it—this is something we all need to master.

Host Dan Shipper: Your mention of this “validation loop” has enormous potential. Beyond automated screenshots and screen sharing, what other hard-core approaches are you exploring?

Mike Krieger:

Our core focus: Can you make it run real workflows—not just inject static data? As systems grow more complex, this becomes harder. For example, we must enable Fable-built iOS apps to log in instantly to our simulation environment, using real test accounts and high-fidelity live data. But we don’t want it to go through the full 8-step tedious new user registration every time it tests a minor button tweak. So we’ve specially developed a high-level permission and encrypted shared key system for AI, allowing it to bypass front-end gates instantly and dive into core business operations—making its test experience nearly pixel-perfect compared to real user experience.

Second: combining known paths with current change paths—known paths are invaluable for regression testing. We’ve expressed ideal workflows in text, which Claude can repeatedly verify. And Claude excels at articulating the intent behind current changes, so this part gets thoroughly exercised. The combination is crucial.

Visual validation is also key—and video is an extremely underused tool for Claude. Recently, I built a prototype: turning Claude’s output into a video recording, using FFmpeg, watching it analyze frame-by-frame, then saying: “This animation is stuttering—I’ll fix it.” Screenshots will never capture that moment—because they miss the instant.

For parts hard to test end-to-end, having Claude build a reliable mock backend—either custom or off-the-shelf—is also fascinating. In the Artifact era, we had comprehensive testing pre-LLM. Each infrastructure piece had a good in-memory implementation for rapid unit testing. Now extending this to Claude’s domain: I’m building something with a robust backend that wouldn’t start on my dev server—it instantly created a great substitute. Over time, this substitute evolves alongside code changes. Previously, I’d say: “Synchronizing this would be a nightmare.” Now I just think: “Claude will read the changes, adapt the substitute, keep both sides synced.” Done.

Host Dan Shipper: There are fascinating architectures: when you receive a bug, an agent automatically fixes it and messages the customer: “Fixed.” Have you noticed any shifts in this flow on Fable?

Mike Krieger:

Several aspects. First, in human-Claude interactions: I keep seeing one pattern: If someone reports a bug in our Slack feedback channel, that thread is passed into the Claude Code session. Thanks to Slack MCP, it can truly pull that thread and respond as me: “This is Mike’s Claude—I fixed it. Here’s the PR link.” Then it adds: “Hold on—still not live. I’ll notify you after deployment.” Hours later: “Deployment went out. Please try it now.” This closed-loop follow-up is relatively new. I’ve had several long-running Claude Code sessions interacting on my behalf. I’ve even included disclaimers in them.

Second, returning to the taste and judgment layer. One aspect is: “A bug reported, so I must fix it.” Another is: sound judgment. Last weekend, I encountered a situation: an internal system had run for ages without reboot, causing a memory leak. Good judgment: “Mike, it’s weekend. Reboot the server now—this solves it immediately. I’ll asynchronously open a PR for long-term fix.” If you want Claude to handle bug-to-fix flow, you truly need it to understand what any good SRE or engineer understands: fix the immediate issue, platform refactoring is a separate matter. Understanding this balance is critical.

What People Should Build With This Model

Host Dan Shipper: What makes this generation of models so electrifying isn’t just raising the floor—enabling anyone with no background to instantly build their own app—but also shattering the ceiling for experts. If you’re a professional engineer or startup founder today, you can single-handedly tackle hard-core projects previously unimaginable. What emerging domains, currently overlooked, could you confidently “go blind” on with this generation of models?

Mike Krieger:

A few ideas—maybe start with fun ones. People constantly have creative ideas about expressing their world’s complexity. Every field has something you deeply understand, and always a version of: “How can I explain this to others? Can I apply tech from elsewhere to my domain?” Take my sister: she’s recently diving deep into environmental engineering—focusing on geothermal energy—filled with mind-bending mathematical models and fluid dynamics simulations. But thanks to Fable’s generational leap in reasoning, she successfully integrated these hardcore, non-specialist technologies into her research. Now, she can even instruct Fable to help her build a full PyTorch-based end-to-end deep learning simulation system—a near-impossible task for a non-computer scientist scholar just a few years ago.

Second: the ability to compose software to solve uniquely personal problems. Internally, we’re heavily investing in making our internal systems as MCP-enabled as possible, paired with proper permission structures and deployment setups. Externally, excellent PaaS platforms exist—just ask Claude, and it’ll set it up. But I especially love that “built something I’ve always wanted” feeling.

Another revelation recently shook me deeply. One of our commercialization team members, non-technical by background, has deeply integrated Claude into every fiber of her daily business workflow. What’s terrifying is she didn’t stop after V1—she quietly iterated intensively for months, behind the scenes, with the model.

This perfectly reveals the most underestimated, yet most seductive, aspect of this generation of reasoning models: In earlier generations, projects often hit a “complexity ceiling.” Once your business logic or code stacked beyond a certain size, models started “losing sight of the bigger picture”—adding new features caused cascading errors, corrupting your existing architecture.

But now, this non-coder colleague, empowered by Fable-level models, has nurtured her system for months in the background. You can clearly see the software growing like a living organism, nourished by AI, evolving relentlessly. Now, she’s rolling out this massive, complex self-built system across our entire commercial department.

An ordinary person with no programmer background can now, single-handedly, push the complexity ceiling of a long-cycle software project to a level that’s literally breath-taking. This is unprecedented in human technological history—a miracle.

Dynamic Workflows

Host Dan Shipper: You mentioned another powerful aspect: dynamic workflows. Can you elaborate?

Mike Krieger:

We frequently prototype such cutting-edge tools internally, and I’m always pressuring the engineers building them: “When will this be publicly released?” Sometimes it’s due to underlying hard infrastructure limitations, but we’re pushing hard to get these gems to market early. For me, dynamic workflows are undeniably among the most globally impressive innovations.

Two key reasons why models like Fable are so powerful. First, they help you scaffold deep, meaningful work. The craziest thing I’ve done: directly handing Fable a complex internal Python project, asking it to fully refactor it into TypeScript—specifically for a concrete deployment consideration.

Back in the Instagram days, executives seriously debated: “Should we rewrite the entire IG codebase in Hack language to seamlessly integrate into Facebook’s infrastructure?” Our conclusion: “Never. It’s not feasible.”

But just last weekend, facing a similarly tangled core codebase, I simply dropped a dynamic workflow into the background and walked away for the weekend. I defined the workflow: deeply understand existing code, generate a spec-like document explaining everything, then translate module by module, perform incremental testing, adversarial validation, check for omissions. When I returned Monday, the miracle happened: it was already a brand-new system running on TypeScript and Bun toolchain—on some architectural levels, even more elegant and faster than my original Python version.

The second, sexier long-term reason: As dynamic workflows spread, in the near future, we can seamlessly distribute subtasks of varying difficulty to model tiers matching their complexity.

Host Dan Shipper: For those who haven’t used it, how did you create that workflow? How did you design it? How do you ensure it’s good?

Mike Krieger:

The entire tuning process is full of hacker-style iterative fun. I started by opening Claude Code and saying: “Bro, I’ve got a monstrous refactoring job—let’s co-design a fully automated workflow.”

It showed me a plan. I said: “Close, but I need three or four additional verification layers to catch missing features.” Then it replied: “Here’s your plan. Ready?” The workflow is code-expressed—this is incredibly valuable; you can see exactly what it plans to do.

After full migration, I had a few minor adjustments. I treated them as mini-workflows, continuing from the previous output. This circles back to the question: Is chat the right interface? Workflows are a sweet spot—use chat to orchestrate, but express them in code, execute in a clean UI, showing every step. I think we’ll increasingly use such approaches to bridge long-range vision with chat.

Curated & Compiled: DeepFlow TechFlow

Disclaimer: Contains third-party opinions, does not constitute financial advice

Recommended Reading

Pharos Network Unlocks AI Model Payment Channels, Introduces New Use Cases for $PROS and USDC as Platform Payment Instruments

06-15
Pharos Network Unlocks AI Model Payment Channels, Introduces New Use Cases for $PROS and USDC as Platform Payment Instruments

NVIDIA Has Plenty of Cash—Why Is It Borrowing $20 Billion?

06-15
NVIDIA Has Plenty of Cash—Why Is It Borrowing $20 Billion?

Will Claude Ban Accounts and Verify ID Cards? Face Recognition Was Old News from Two Months Ago, and "Handing Over Data to Police" Is a Misinterpretation

06-15
Will Claude Ban Accounts and Verify ID Cards? Face Recognition Was Old News from Two Months Ago, and "Handing Over Data to Police" Is a Misinterpretation

Japan's Central Bank on the Brink of Rate Hike—Can the AI Bull Run Withstand?

06-15
Japan's Central Bank on the Brink of Rate Hike—Can the AI Bull Run Withstand?

5-Second Breakthrough with Just 1 Interaction: Has the "Strongest Security Mechanism" of Claude Fable 5 Been Cracked by a Chinese Team?

06-13
5-Second Breakthrough with Just 1 Interaction: Has the "Strongest Security Mechanism" of Claude Fable 5 Been Cracked by a Chinese Team?

Why Is the "AI Service Subscription Model" Inevitably Headed for Extinction?

06-13
Why Is the "AI Service Subscription Model" Inevitably Headed for Extinction?

Managing a company valued at nearly a trillion dollars, Anthropic's CEO has only one direct report.

06-13
Managing a company valued at nearly a trillion dollars, Anthropic's CEO has only one direct report.