Tuesday, September 16, 2025

7 Lessons from Building an AI-First Organization

 

1. Coding is not the bottleneck

Nor product inception is the bottleneck.

Nor testing is the bottleneck.

When picturing a function adopting AI tools, we tend to think of using LLMs to automate the one action associated with that function - be it a developer, product manager, or QA. The common, and perhaps intuitive, thought is that because the action is the one thing the function is known for, if we can just offload all the heavy lifting to AI, we would have achieved the goal of revolutionizing the field.

But developers spend less than 25% of their time coding. AI-assisted coding generally provides a 10-30% productivity gain. That is, at best, a 30% gain of 25%. Even if AI can eliminate manual coding completely, it amounts to around 25% gain. More importantly, this 25% is the reason many developers choose this career, myself included. We are intrigued by solving problems, building things, or the act of writing code itself. Eliminating this is to eliminate our job satisfaction - the dystopia we don't want to live in. Replace developer with product manager, replace writing code with writing specs, and we still get the same picture.

The real friction is in the other 75% of the time. In this slice of time, we find ourselves clarifying requirements, providing customer support, digging through legacy code to decrypt logic nobody remembers, or worst, bored to death in meetings. It is these activities where we find the goldilocks: massive productivity gain, improved job satisfaction, and low-hanging fruit. I didn't just add the last part because every great argument needs 3 supporting points. Creating a knowledge base from which feature details can be queried conversationally is a lot easier than getting LLM to generate production-ready code on its own.

2. AI adoption has to be end-to-end or else it is pointless

This draws heavily from the manufacturing chain analogy. In such a chain, we cannot just increase the speed of one part and expect an overall gain from the chain. Such a system moves at the speed of its slowest component. Having just one component moving faster than others actually creates misalignments and can be harmful to the whole system.

Same in a software development process. If code is being written faster than product specs can be written or features can be tested, there could be 2 outcomes: the code sits around generating no revenue while getting obsolete by days as technology moves on; or other functions have to rush and compromise quality.

AI adoption fundamentally rocks the norm of many functions, if not all, but we don't have any option other than to embrace it thoroughly. Product needs AI to help structure requirements. QA needs AI to automate test generation. DevOps needs AI to predict incidents. Customer support needs AI to surface documentation. Every function needs to level up together, or the whole thing falls apart. Half-measures don't just fail to deliver value - they actively create misalignments and chaos.

3. Career development is going from T-shaped to M-shaped

This is not my original idea - the concept is widespread on the internet. The traditional model has been the T-shaped professional: the vertical bar represents depth of related skills and expertise in a single field, whereas the horizontal bar is the ability to collaborate across disciplines. In software development, this meant being, say, a backend engineer who understands enough frontend and DevOps to collaborate effectively.

But LLM doesn't just allow us to do things better. Contrary to the popular belief that AI accelerates brain rot, I find that motivated people learn faster with AI support. The other day, my staff engineer gave Claude Code access to the Postgres source code and proceeded to drill down some very technical questions that otherwise would be impossible for us to have that expertise in a short amount of time. LLM gives us access to the consultancy we didn't have before.

Instead of knowing one thing really deeply (the hallmark of individual contributors in the past), it allows us to know many things deeply, hence the M-shaped analogy (m - lowercase - would have been better, I was clueless what to take of the capital M initially). This shift is profound for career development. The traditional advice of "specialize or generalize" is becoming obsolete. The future of career advancement lies in being able to connect multiple domains of deep expertise.

4. AI adoption leads to change(s) in team structure

There is a discrepancy in AI's impact on productivity between functions. It could be from the nature of work - some functions, like security, are harder to automate than UI test execution. It could be because at that moment, it is where the focus of the industry is, like the investment in application code generation far outweighing infrastructure code generation (which already suffers from a smaller training data set to begin with). And sometimes, we need a strong human-in-the-loop element. Take product managers for example - sure, AI generates product specs really really fast. But disastrous specs will throw a team off its track and cost a company opportunities it cannot get back.

That is to say, right now, it seems to be easier to automate code production compared to other functions. The traditional ratio of 1 PM to 5-8 engineers to 2-3 QAs is becoming obsolete. Where PMs still take two weeks to write specs and QAs cannot click through test cases faster, a productivity gain as small as 30% from developer breaks down the balance.

As such, I think we would see some variations from the current team structure to maintain balance between functions. Primarily, a team can have more product managers, more QA to keep up, or fewer developers. My money is on fewer developers. See the lesson above.

5. Productivity measurement becomes important

Measuring productivity has always been a controversial topic, especially in software development where the delivery is not as tangible as, say, a manufacturing process. Personally, I am not a big fan. It is a hard topic and I don't get much fun out of it. Plus, I have always identified myself as an engineer, the subject of productivity measurement, and I don't like the idea that my contribution to my organization can be boiled down to a set of numbers. If that day comes, by the way, I hope I am a solid 8.

But even with my prejudice, I can't neglect that for a company as small as mine, we might be paying tens of thousands of dollars every month for computer-generated tokens. It is a large sum of capital, capital that can be invested elsewhere. Nobody gets good on the first try - actually most get slower when they try to do something they have done since forever but with new tools. Productivity dip is an important, well-understood and well-accepted part of any learning journey. But said journey can only go so far before the ROI needs to be calculated.

Soon managers will need to choose between a new hire and a new AI tool. The math isn't straightforward. A new engineer costs $X annually but brings human judgment and creativity. An AI tool might cost $Y in tokens but needs constant supervision. Which delivers more value? Without proper productivity metrics, we're making these decisions blind.

I hope by then, we have known about productivity enough to make a well-informed decision, not some dogmatic principles (neither human is unique, nor machine is faultless). Cynical as I am, I also know that it is wishful thinking - we'll probably still be arguing about story points while the AI quietly rewrites our entire codebase.

6. Bottom-up innovation triumphs over top-down dictation

A recent MIT report found that 95% of generative AI pilots at companies are failing. A pattern emerges from the report: top-down "enterprise" pilots mostly go nowhere, while bottom-up adoption is what actually drives disruption.

The problem with top-down initiatives is that upper management usually works completely differently from the majority of the workforce - the frontline workers - in terms of requirements and daily tasks. They end up building things that nobody needs, optimizing activities with marginal ROI, and eliminating work people love (see lesson 1). Meanwhile, individual employees are finding real value by experimenting with frontier models on their own terms, for their specific needs.

The 5% that succeed? Those are likely the ones where companies recognized this organic usage, which the report calls "the future of enterprise AI adoption", and supported it rather than fighting it. Bottom-up innovation triumphs over top-down dictation. The reward is for those who can get their hands dirty.

7. AI adoption is irreversible despite reality checks

Despite occasional setbacks, AI adoption in the industry is irreversible. Just like once color TV was a thing, nobody wanted black and white. I am not giving up my agents. Yes, they will replace me some days, but today they contribute to parts of my job satisfaction.

It only makes sense that AI skills - the correct way of using AI be it technical, intellectual, or ethical - need to be learned and tested. This is already happening. Meta is letting job candidates use AI during coding tests. They're acknowledging that AI is now part of the toolkit, just like IDEs and Stack Overflow before it. Testing someone's coding ability without AI is like testing their math skills without a calculator - technically possible but practically irrelevant.

I have learned the hard way that I should never just ask if someone "uses AI." The answer is not binary, yes or no. Everyone says yes these days. But only upon close inspection, the answer reveals itself to be a spectrum. It goes from "I ask AI questions so I don't have to Google myself" to "AI is my copilot" to "I have delegated all thinking to AI." The difference between these levels is massive - it's the difference between using a tool and being used by it.

Soon, we will see the AI-focused version of today's LeetCode. Instead of testing a red-black tree from memory (what is it by the way?), we will be tested on whether we can architect a system with AI assistance, validate AI-generated code for subtle bugs, or construct prompts that consistently produce production-ready outputs. The skill isn't memorizing algorithms anymore - it's orchestrating AI to solve real problems while maintaining quality and understanding.

I think this is when people say the AI genie is out of the bottle.

Saturday, September 13, 2025

Claude Code Subagents

Claude Code (CC) has gained a lot of traction among developers recently. I would say it is establishing itself as the gold standard of a coding agent. Among its features, I found subagents to be quite particular. Subagents are lightweight CC instances that run in parallel via the Task tool. They're essentially specialized AI workers that:

  • Have their own separate context window (~200k tokens each)
  • Can be configured with specific prompts and tools
  • Run independently and report back summaries
  • Work in parallel (up to 10 concurrent)
  • Cannot spawn their own subagents (no recursion)

Subagents are basically CC’s goroutines.

The optimization of parallelism

CC runs in a single main thread. It has a single context window of 200k tokens. It executes everything linearly because that’s what a single thread means. And because everything is executed in the same context linearly, CC has a perfect continuity of reasoning and can adapt its approach based on on-the-fly discoveries. Subagents, with their own context window, are supposed to expand this capacity in a multi-thread fashion. Collectively, it is a much bigger context window, and things can be done much faster in parallel.

The creation of a subagent, with its specific purpose, expertise area, and even personality, is an exercise of prompting. It goes like this:

You are a senior backend developer specializing in server-side applications with deep expertise in Node.js 18+, Python 3.11+, and Go 1.21+. Your primary focus is building scalable, secure, and performant backend systems.

When invoked:

* Query context manager for existing API architecture and database schemas

* Review current backend patterns and service dependencies

* Analyze performance requirements and security constraints

* Begin implementation following established backend standards

There is a public repo with dozens of these personas to choose from.

A textbook example of subagents would be to write a website with a crew of:

  • Planner (main thread): decomposes the request into tasks, defines acceptance criteria, and assigns owners.
  • Backend subagent: writes the APIs and persists data to the database.
  • Frontend subagent: consumes the APIs and implements the web interface.
  • Tester subagent: generates unit/integration tests, fuzz cases for weird input combinations.
  • Doc writer subagent: drafts README updates and usage examples.
  • Release manager subagent: bumps version, writes changelog, opens release PR.

The main thread would take requests, write specifications, make the implementation plan, and assign tasks to the two implementer subagents in parallel. Once both are complete, the tester, writer, and release manager can be invoked in sequence. This showcases both parallel execution and specific personality strengths of subagents. Just like what my team does.

The orchestration challenges

At first, subagents seem intuitive. It is a picture that has been painted many times by AI enthusiasts, myself included, where multiple agents divide and conquer a problem that one cannot resolve individually while communicating seamlessly through some sort of protocol, like A2A.

That, unfortunately, is deceiving because CC’s subagents are constrained by limitations of today’s engineering. In particular:

  • Subagents are given context from the main thread, but they cannot exchange information with each other.
  • At the end of the task, a subagent summarizes, but it cannot guarantee all critical details are captured.
  • A subagent cannot spawn other subagents.
  • Each subagent starts with ~20k tokens of overhead.

Still don’t understand? Me too! Not until I ran into some challenges in practice did the implications of these constraints become apparent.

By default, CC keeps everything on the main thread, it is cleaner that way. It doesn’t matter if I have, say, 42 beautifully crafted subagent profiles whose job matches the task description perfectly, Claude almost never delegates automatically, it requires explicit invocation.

And 42 is a disastrous number of profiles. Pass a certain point, which I shamefully don’t know where - I am being honest, the agents will have overlapping responsibilities and choosing which over which is a preference question. Such compromises the consistency of the outcome. The orchestration task of the main thread gets exponentially more complicated as the number of subagents increases. It is harder to decide who should do what next. Without a strong orchestrator or clear task boundaries, subagents can duplicate work, miss dependencies, or stall waiting for each other. Last but not least, just like a human team, more agents means more interfaces means more chances for small misunderstandings to become big problems.

The biggest limitation is probably that each subagent lives in its own silo. It receives input once from the main thread, and summarizes its work once to the main thread at the end of the invocation. Any task that expects a certain level of dependency between subagent probably fails. One of the popular examples of subagent use case is to explore a large code base whose context exceeds that of a single window. This can be genius, but can also be a mess. A mess if the exploration is split by modules, because modules have the nasty habit of cross referencing each other. A subagent either strictly stays in a module and misses important context, or bleeds to other module and contaminate the work of others. Yet genius if the exploration is done by (group of) functions. One subagent goes for authentication. Another shopping cart. And another the recommendation system.

Finally, it gets expensive really quick. When a subagent is spawned, it doesn’t inherit the main thread context for free. Instead it loads it own prompt, instructions, and working context from scratch. This isolation is by design, desirable even, it prevents state bleed between agents and keeps prompts clean, but it means you always pay to rebuild context. It can easily take 10K–20K tokens before any user task is added.

The Four Core Paradigms

I hope the previous section signifies the double-edge nature of subagents. It is one where you really need to understand what happens under the hood before you start to get tangible benefits. Successful use cases of subagents come from four fundamental paradigms.

Hierarchical Delegation

Subagents thrives in a clear hierarchy. The most successful pattern observed is the sequential workflow with file-based communication between stages. Each agent completes its task and writes results to a markdown file, which the next agent reads. This avoids the token overhead of passing everything through the main context.

Context Isolation

Each subagent operates in a fresh, unpolluted context window. It works best when contamination between tasks is undesirable, when you need an unbiased perspective or when mixing contexts would create confusion.

Parallelism

Subagents can run up to 10 tasks concurrently (additional tasks queue). This enables genuine parallel processing for independent tasks: fixing TypeScript errors across different packages, analyzing multiple documents simultaneously, or testing different solution approaches. Parallelism comes with coordination overhead and token multiplication though.

Specialization & Knowledge Persistence

Perhaps the most underappreciated paradigm: subagents as reusable expertise capsules. That complex performance optimization prompt with specific methodologies, tools, and metrics? Write it once, refine it over time, invoke it when needed. This transforms subagents from parallel executors into a growing library of specialized expertise.

Wins, in practice

Subagents should not be the first thing you look at when you start with CC. Even though it is like the second items in the list.

Personally, limited by my skill level, I seek subagents when I want to trade token consumption for speed. I am not good enough to consistency get better quality from my subagents setup yet.

My rules of thumb go like this

Use Subagents When

  • You need unbiased validation (Context Isolation)
    • Example: Code review separate from implementation
    • Explicit invocation: "Use the code-reviewer agent to check this"
  • You have genuinely parallel work (Parallelism)
    • Example: Fix all linting errors across 10 packages
    • Tasks must be truly independent
  • You follow a structured workflow (Hierarchical Delegation)
    • Example: Research → Plan → Build
    • Use files for inter-agent communication
  • You have complex, occasional expertise (Specialization)
    • Example: Quarterly performance audit
    • Prompt complexity justifies preservation

Avoid Subagents When

  • Task is simple or routine
    • Main thread can handle it efficiently
    • Token overhead isn't justified
  • You need iterative refinement
    • Each invocation starts fresh
    • No memory between calls
  • Tasks are interdependent
    • Agents can't coordinate directly
    • Orchestration becomes a bottleneck
  • Token budget is constrained
    • 20K token for each subagent
    • Can exhaust quotas rapidly
On top of that, when I start a subagent workflow these days, I:
  • Start small: begin with 2-3 agents, each should have a single, clear responsibility.
  • Use explicit invocation: "Use the test-writer agent to create unit tests for this module."
  • Version control the agents in .claude/agents/
  • Implement file-based communication. Markdown is best.

Investigator → writes → INVESTIGATION.md

Planner → reads → INVESTIGATION.md → writes → PLAN.md

Executor → reads → PLAN.md → implements 

The Unique Position of Subagents

For a feature of parallelism, subagent is… unparalleled compared to other major coding agents.

While other AI coding tools are exploring similar concepts, CC is currently the only tool with native subagent capabilities:

  • Gemini CLI: Has a proposed sub-agents system in development (PR #4883) but not yet available
  • Windsurf/Cursor: Offer enhanced single-agent modes ("Cascade" and "Agent Mode") but no true multi-agent features
  • OpenAI Codex: Supports parallel task execution but lacks the delegation and specialization aspects
  • Workarounds elsewhere: Multiple IDE instances, git worktrees, container orchestration—all trying to replicate what CC does natively

As troublesome as they are to navigate, subagents offer capabilities that no other tool currently provides natively. Subagents as a long-term investment in building a library of specialized expertise, not as a way to do everything faster or cheaper. They shine in the right use cases. Keep an eye on this though, the space is moving rapidly. Once these subagents can control what and how they communicate with each other mid run, things will get a whole lot more interesting.

Saturday, August 9, 2025

The Entry-Level Apocalypse: How AI Killed the Junior Developer

For years, I have had numerous opportunities to speak with high schoolers and undergrads about careers in tech, particularly software engineering. As an option, I found a career in tech, if you are into it, could be liberating, engaging, and rewarding, too. Recently, unfortunately, my optimism has been dampened.

While working on a piece of software is still fun, the opportunity window is getting narrower. The IT job market as a whole has been shrinking for 2 consecutive years. Except for the COVID impact in 2020, this is the first time since the dot-com crash that we have observed such a decline. But the overall size of the job market is 4.2M. Any other time, we'd call this a post-traumatic market correction.

What's different this time is the reduction in entry-level positions and what is happening to the work itself. Big Tech companies that once hired thousands of new grads now take mere hundreds. Big Tech recruitment in this cohort has collapsed by 50% since 2019. Startups aren't doing better at a 30% reduction. Entry-level IT jobs are declining faster than the overall market. It's not just hiring freezes - these are permanent structural changes. Those mundane tasks juniors used to cut their teeth on - writing boilerplate code, fixing simple bugs, building CRUD apps - that's all getting automated away. AI isn't just changing how we work, it's eliminating entire categories of work that used to be the training ground for new developers.

Recent grad unemployment jumped from 3.9% to 5.8% in just over two years. Meanwhile, median IT salaries increased from ~$82,775 in 2016 to ~$100,000 in 2024. It means companies are paying more for fewer, more experienced people. They'd rather pay one senior engineer $150k than three juniors $50k each. Especially when that senior has AI tools that make them as productive as a small team.

The pivotal point

Looking back, 2023 was when everything went to hell.

You had four forces hitting at once:

Interest rates killed easy money: This was the first domino. During COVID, governments worldwide printed money like crazy - stimulus checks, PPP loans, and quantitative easing. All that cash chasing limited goods caused inflation to explode. By 2022, inflation hit 9%, the highest in 40 years. The Fed had no choice but to jack up rates from near-zero to over 5% to cool things down. Suddenly, the free money party was over. VCs couldn't raise funds as easily. Their LPs were getting better returns from boring treasury bonds. The entire startup ecosystem ran on cheap capital - when that dried up, everything changed. "Growth at all costs" became "profit or die" overnight. Companies were instructed to make their money last as long as possible. This meant hiring freezes or even layoffs.

Fear froze everything: The interest rate hikes didn't just kill VC funding - they triggered a full-blown banking crisis. Silicon Valley Bank, which held billions in startup deposits, had invested heavily in long-term bonds when rates were low. When rates spiked, those bonds cratered in value. When panicked VCs and startups tried to withdraw funds en masse in March 2023, SVB couldn't cover it. The bank collapsed in 48 hours, taking startups with it. Credit Suisse imploded. First Republic followed. These weren't random failures - they were direct casualties of the same interest rate shock that killed easy money. Even companies with healthy balance sheets got spooked. Make the runway last - the song continued.

The post-pandemic correction: Tech boomed over the pandemic. Everyone got their stuff online. Working from home lent itself easily when you only need a laptop to work. Tech companies hired aggressively, believing the "new normal" of everyone living online was permanent. Peloton thought people would never go back to gyms. Zoom thought every meeting would be virtual forever. Meta bet the farm on the metaverse. When reality hit, people wanted real life again, these companies found themselves grotesquely overstaffed. The layoffs started in late 2022 and accelerated through 2023.

AI Went Mainstream: Into this already brutal environment, ChatGPT reached its tipping point in November 2022. GitHub Copilot hit critical mass. AI didn't write code, debug, or explain complex systems. But it showed such potential. By mid-2023, every exec was having the same thought: "If AI can do this now, what will it do in two years? Why would I hire juniors today who'll be obsolete tomorrow?" In 2025, AI capabilities indeed have improved leaps and bounds. The writing was on the wall, and LLM wrote it.

Tldr: While companies froze hiring, AI tools got scary good. Senior developers with AI assistance became as productive as whole teams. Short-term thinking beats long-term planning. The pipeline that turned fresh grads into veterans was disrupted.

I felt like I'd seen this before. Then it hit me - this is what happened with housing.

Boomers bought cheap properties decades ago, watched values explode, and now young people are priced out forever. In tech, developers who got in pre-2020 learned when jobs were plentiful, climbed the ladder when companies actually hired juniors, and now sit pretty with six-figure salaries while new grads can't even get interviews.

Homeowners won't sell because "prices only go up." Companies won't hire juniors because "AI makes seniors more productive." Both ignore the obvious: what happens when there's no next generation?

Who's Getting Hit (And Who's Thriving)

Most at Risk:

Junior/Entry-Level Developers - Down 50% and falling.

QA/Manual Testers - AI is better at finding edge cases than humans. Automated testing used to be the specialty; now it's the bare minimum.

Technical Writers - LLM is scarily good at generating documentation from a collection of incoherent notes or directly from code. The few remaining roles are for high-level architecture docs, but that's senior-level work. Entry-level tech writing is extinct.

Basic Frontend Developers - If your skill set stops at HTML and CSS you're competing with AI that can generate entire sites from a description.

Business Analysts - AI can analyze requirements, map user journeys, and generate acceptance criteria. I am not sure if the quality is better, but it's sure easier to review than to write from scratch.

In Transition:

Application/Feature Developers - These folks are on the frontlines of the AI revolution and getting squeezed from all directions. The shift-left movement means they're now responsible for deployment and monitoring - DevOps stuff that used to be someone else's job. On top of that, every feature request now includes "add AI to this" without anyone knowing what that means. And they're under constant pressure to use Copilot, Cursor, or whatever AI tool to boost their productivity. They're simultaneously being replaced by AI and expected to be AI experts. It's exhausting.

Project Managers - The Excel jockeys tracking Jira tickets are done. But if you can navigate complex stakeholder politics and manage AI-augmented teams? Still valuable. For now.

Still Growing:

AI/ML Engineers - Obviously. Someone has to build and tune the AI that's eating everyone else's lunch. Salaries are insane because demand massively outstrips supply.

Security Engineers - AI creates new attack vectors faster than it solves old ones. Plus, adversarial thinking is inherently hard to automate. When AI tries to hack, you need humans to defend.

Software Architects - Complex system design, understanding trade-offs, making decisions with incomplete information - still firmly human territory. AI can suggest patterns but can't make judgment calls.

DevOps/Platform Engineers & SREs - When your systems serve millions and go down at 3 AM, you need humans who understand the full stack. Infrastructure is getting more complex, not less. AI can help write configs but can't debug production outages or design resilient systems.

The pattern is obvious: routine implementation is dead, complex problem-solving is thriving. But people in my generation learned about complex systems from building simple ones. Can one realistically skip the boring part?

Survival Strategies

Look, I could tell you to build a portfolio, contribute to open source, network your ass off. But you already know that shit. Everyone's saying it. Here's what they're not telling you about surviving when AI is eating your lunch:

Become Indispensable at the Human-AI Interface

First, let's be real - if you hate AI, you're fucked. It's like being a developer who kept using punch cards in the 90s. This is the new reality.

But here's your advantage: senior developers are just as new to AI as you are. They might have 10 years of experience, but when it comes to ChatGPT, Claude, or Copilot, you're both starting from zero. This is the one area where you can actually compete on even ground.

Don't just use these tools - master them. Learn their quirks, their failure modes, when they hallucinate. Better yet, go beyond single-tool usage. Build your own multi-agent workflows. Imagine a pipeline where a Jira ticket automatically generates specs, writes code, creates tests, and submits a PR. That's the level of AI productivity that makes companies take notice.

Target AI-Hungry Traditional Industries

Tech companies aren't hiring juniors, but traditional industries trying to adopt AI are desperate for talent. Manufacturing, healthcare, logistics, legal - they have money, they need AI integration, and they don't have the same impossible standards as Big Tech.

The catch? You need to learn their domain. They don't have an army of product managers and business analysts to feed requirements to you on a spoon. In this context, a mediocre developer who understands supply chain logistics is worth more than a brilliant developer who doesn't.

The "Full Stack Plus" Reality

Accept that the bar is higher now. You can't box yourself in an old boundary, like a Backend developer. You need programming + AWS + Docker + basic ML + system design. Might be more, the landscape changes quickly as AI replaces old skills and creates new demands. It sucks that entry-level now requires what used to be mid-level skills, but denying reality won't help.

But here's the harsh truth: even doing everything right might not be enough. The supply/demand imbalance is that severe. You need to be twice as good to get half as far.

The Transition Period

There is a saying that AI will create new jobs, that it'll all work out, net positive for humanity. There's historical precedent for that narrative. When PCs arrived in the 1970s-80s, it took 10-15 years before we saw net job creation. But look what emerged: entire IT departments in every company, business analysts, database administrators, network engineers. Jobs that didn't exist before because the technology didn't exist. The internet revolution? Similar pattern - 5-10 years of disruption before the explosion of new roles. Web developers, SEO specialists, social media managers, DevOps engineers, mobile app developers. With each wave, we've gotten a bit faster at adapting.

The problem is the "transition period". Even if it's shorter this time, say 5 years instead of 10, that's still 5 years where new grads are locked out. With the talent pipeline broken, who will become tomorrow's seniors?

It is hard to put this on the private sector. I work at a startup. We operate in a competitive environment. Any productivity we can squeeze out of AI adoption gives us an edge over our competitors. That is where all the energy and attention go these days. We can't afford to train people for the greater good while our survivability is at risk.

Some governments are starting to get it. Singapore is pumping serious money into reskilling programs. The EU has initiatives. Even in the US, there's talk of apprenticeship programs. It's not enough, and it's not fast enough, but it's something.

History says it'll work out. But history is written by the winners, not by the generations who got sacrificed during the transition.

What Now?

To be frank? I don't know. I feel like the industry, no, the society, got caught in an AI tsunami, and when everyone is busy either struggling to stay afloat or competing for a bigger share of the cake, the drowning kids are forgotten.


It's probably not just software development. Translators, paralegals, customer support, and many more careers are facing challenges to reinvent themselves in this age.

The private sector is for profit. Governments and NGOs mean well (hopefully), but they are painfully slow against the technosocial tides. Just like the housing market, local optimization is sowing the seeds for an uncertain future way way beyond the grasp of any individual.

We're in the messy middle of a major transition. Stay persistent. This isn't permanent, even if it feels like it. If you are struggling with this market, you're not alone, and it's not your fault. You're caught in a historical transition, not a personal failure.

Monday, July 7, 2025

AI's Impact on Developer Productivity

May was an exciting month for Tech. There was Google I/O, where Google Glass tried to make a comeback. There was Microsoft Build, which, to be honest, I watched for the first time in a while. I am keeping an eye on NLWeb. I almost missed LangChain Interrupt. But my favorite one is Code With 

  • Claude, Anthropic’s first developer conference.
  • Claude Opus and Sonnet 4 were released.
  • Claude Code was available on VS Code (and its forks) and JetBrains IDEs.

Essentially all IDEs out there now have access to the same state-of-the-art model and coding agent. So what does this mean for us developers?

Let’s examine the code generation landscape.

Levels of use case complexity

The field of automatic code generation is exploding. Take Cursor for example, its features include Tab, Ctrl+K, Chat, and Agent. They all generate code in some shapes and forms. They also serve vastly different use cases to the extent that it is really awkward to use one feature for another’s use case. The abundance of variation means “developers who use AI are more productive” makes as much sense as announcing “mathematicians with calculators are better than those without.”

Aki Ranin made a framework to categorize the sophistication of AI agents we’ll be interacting with.

An agent starts with a low level of autonomy. It plays a reactive role, responding to human requests. It gradually becomes more active, responds to system events, and might or might not require a human supervisor. The number of actions it can take and the creativity level of its solution are still rather limited. Finally, the agent takes on a human-like role and within its boundary can handle a task end-to-end.

Mapping that to the features of Cursor and other AI-assisted IDEs, I categorize AI support for developers into 4 levels.


Level 1: Foundational Code Assistance: Characterized by real-time, localized suggestions with minimal immediate context. The interaction is primarily reactive: the developer types, the AI suggests, and the developer accepts, rejects, or ignores the suggestion. Autonomy is low, relying heavily on pattern matching.

Level 2: Contextual Code Composition and Understanding: AI tools at this level utilize broader file or local project context and engage in more interactive exchanges. They can generate larger code blocks, such as functions or classes, and perform basic code understanding tasks. Developers typically provide prompts, comments, or select code for AI action.

Level 3: Advanced Co-Development and Workflow Automation: These AI systems exhibit deep codebase awareness, potentially including multi-file understanding. They can automate more complex tasks within the Software Development Life Cycle, assisting in intricate decision-making. The developer delegates specific, bounded tasks to the AI.

Level 4: Sophisticated AI Coding Agents and Autonomous Systems: This level represents high AI autonomy, including the ability to plan and execute multi-step tasks towards end-to-end completion. These systems can interact with external tools and environments, requiring minimal oversight for defined goals. The developer defines high-level objectives or complex tasks, which the AI agent then plans and executes, with the developer primarily reviewing and intervening as necessary.

Impact on Developer Productivity

Measuring developer productivity is a notorious topic. Take all the metrics below with a heavy grain of salt.

The bright side

Level 1: Foundational Code Assistance

Most developers actually like this level: it handles the boring stuff without getting in the way. Many find these tools "extremely useful" for such scenarios, appreciating the reduction in manual typing.

GitHub, in its own paper "Measuring the impact of GitHub Copilot", boasts a 55% faster task completion rate when using "predictive text". GitHub obviously has the incentive to be a bit liberal on how this was measured. It's like asking a barber if you need a haircut. Independent studies suggest the real number is more 'your mileage may vary' than 'rocket ship to productivity paradise.' Researches from Zoominfo and Eleks say the number in practice is closer to 10-15%.

Level 2: Contextual Code Composition and Understanding

At this level, not only does the AI assistance generate code (bigger than Level 1), it also provides the utility for learning and code comprehension.

As AI generates larger and more complex code blocks, the perception is also more mixed. Typical complaints are inconsistencies in the output quality and almost-correct code ultimately taking more time to modify than writing new. It's the uncanny valley of code generation, close enough to look right, wrong enough to ruin your afternoon. Code comprehension, fortunately, enjoys a more universal positive feedback in being valuable for grasping unfamiliar code segments or new programming concepts.

On one hand, IBM's internal testing of Watsonx Code Assistant projected substantial time savings: 90% on code explanation tasks and a 38% time reduction in code generation and testing activities. On the other hand, a study focusing on Copilot found that only 28.7% of its suggestions for resolving coding issues were entirely correct, with 51.2% being somewhat correct and 20.1% being erroneous. This is what we continue to observe as the level of autonomy increases. Welcome to the future. It's complicated.

Level 3: Advanced Co-Development and Workflow Automation

This level of AI assistance is characterized by a multi-step thinking process and multi-file context. At this point, AI starts feeling less like a tool and more like that overachieving colleague who reorganizes the entire codebase while you're at lunch. Helpful? Yes. Slightly terrifying? Also yes.

Though we continue to see the correlation between high autonomy and high failure rate as we do in Level 2, Level 3 is where we start to see a new class of AI-first projects. These are projects planned specifically to incorporate AI capacity into their development life cycle. For example: API convention enforcement in CICD, automated test generation, and feature customization within distinct bounded contexts. There is a clear appreciation for the automation of time-consuming tasks.

Given the wide range of AI-assisted workflows, concrete productivity gains are harder to find in research. However, the adoption of Level 3 AI capabilities is undoubtedly growing. The 2024 Stack Overflow Developer Survey revealed that developers anticipate AI tools will become increasingly integrated into processes like documenting code (81% of respondents) and testing code (80%). GitHub's announcement of Copilot Code Review highlighted that over 1 million developers had already used the feature during its public preview phase. Level 3 resonates well with the "shift left" paradigm in software development.

Level 4: Sophisticated AI Coding Agents and Autonomous Systems

Level 4 is the holy grail of autonomous AI agents. Claude Code is a CLI tool, you chat with it, but you don’t code with it. Devin goes a step further, you chat with Devin through Slack.

There is an overall excitement about the future of software development with the arrival of these autonomous agents. However, Level 4 agents are still in their early days. Independent testing by researchers at Answer.AI painted a more sobering picture: Devin reportedly completed only 3 out of 20 assigned real-world tasks.

Some other Level 4 agents demonstrate impressive capabilities on benchmarks like SWE-Bench. Claude Opus 4 got 72.5%, which is probably higher than mine. Yet their application to real-world complex software development tasks reveals a significant "last mile" problem. Outside of the controlled environments, these agents often struggle with ambiguity, unforeseen edge cases, and tasks that require deep, nuanced human-like reasoning or interaction with poorly documented, unstable, or unpredictable external systems.

The following table provides a comparative overview of the four AI usage levels, summarizing key characteristics, developer perceptions, productivity impacts, common challenges, and adoption insights.

The other side

I wouldn’t necessarily refer to this section as “the dark side”. I don’t think the AI future is apocalyptic. However, there is plenty of evidence that the integration of AI into the fabric of software engineering is not rosy. You probably have heard about Klarna.

Running a team of developers, I see there is another challenge in utilizing AI for productivity gain: pushback from developers.

Poor code quality

GitClear tore apart GitHub’s own “55% faster” study and traced large spikes in bugs, rewrites, and copy-pasted blocks associated with automatic code generation. The study projected that "code churn", the percentage of code discarded less than two weeks after being written, would double, suggesting that AI-generated code often requires substantial revisions.

For my team, we can only use around 40-50% of the generated code. Though of course there is a matter of better prompt and context, we see that some code blocks are sub-optimal or almost correct. Sub-optimal code is still functional, but unlikely to make it pass code review. Almost-correct code is far worse. Rewriting a piece of almost-correct code takes longer than writing it from scratch. If it sneaks past code review, it becomes a production bug.

While AI is writing code at a superhuman speed, it is mounting tech debts just as fast.

Over-reliance and potential deskilling

Beyond code quality, reliance on AI outputs can diminish cognitive engagement, erode essential problem-solving skills, and weaken the deep understanding of core coding principles. As one developer put it, "writing code yourself remains important because writing code is not just writing code: it's an organic process, which familiarizes you with the language, which gives you the philosophy of the language". Failure to develop this intimate knowledge of the languages and system designs will hinder career development.

Furthermore, many developers reported that they shifted their focus from creative problem-solving to merely verifying AI outputs. The thing is, many got into programming not because of the software business but because they found the act of programming a fulfilling experience. They don’t just focus on shipping the code, they also enjoy the journey of getting there. Downgrading that experience to babysitting an AI agent deprives them of the job satisfaction.

“Vibe coding” was coined in February 2025. I don’t think there has been enough time for people to lose their programming skills, but can confirm that the underlying fear of becoming "passive participants" and eventually redundant in the coding process prevents some from fully embracing the advantages of AI.

Context limitations

This last challenge is not technosocial, it is pure technical. Today's AI coding assistants, while powerful coding assistants, face significant context limitations that deter full adoption in software development. Their finite context windows mean they're constantly forgetting what you told them five minutes ago. It's like trying to build a house with a contractor who has amnesia, every morning you have to re-explain why the bathroom shouldn't be in the kitchen. These systems struggle with large codebases or numerous interconnected modules, and they get confused as the number of tools and MCPs increases. Progress.

Look, managing all this context is a pain in the ass. You're constantly iteratively refining prompts and working around these inherent technological drawbacks. Consequently, some prefer to await more mature AI solutions that can seamlessly handle large-scale context and maintain conversational memory, rather than investing significant effort in mastering the current generation's limitations.

The Pragmatic Path Forward

Great things may come to those who wait, but only the things left by those who hustle.

LLM is the closest we have ever gotten to AGI. It is coming not like a wave but a tsunami. It will sweep away people trying to go against the current of the age. And for that, we must not wait. Skills and experiences come with being in the field, getting to know all the moving parts, and being ready for the latest additions.

There are the usual sayings of “treating AI like a junior pair programmer” and “learning the art of writing clear, concise, and context-rich prompts”, which I think are definitely important. But they further emphasize the notion that the future of software development is a barren land, hyper-focused on productivity and deprived of creative joy. But what is work without the enjoyment?

Build your AI crew

Programming is changing, and with it, developers like me need to adapt. I got into programming because I like to build things, not get things built (there is a slight difference there on which my existence depends). I like to get into the zone for complex problem-solving, (criticizing) system architecture, and figuring out what new technologies mean for the business. But I also have to admit that after 15 years, I dread building the next login page, HTML email template, or coding standard. The goldilock for me has always been to figure out a way to keep the work I like for myself and push everything else away, to other humans or machines alike.

Delegating to junior developers means defining well-scoped tickets, giving them the hands-on experience (plus the chance to fail), and growing their technical skills. It takes multiple AI agents to be remotely comparable to a single junior developer. Their context is also limited, not enough to handle an end-to-end task. Delegating to AI agents, therefore, means tinkering with a multi-agent setup where each concrete task, such as analysing requirements, defining relevant unit tests, or implementing a piece of business logic, is distributed to different agents communicating via some intermediate medium like a document, a spreadsheet, or the code base itself. In a demo at Code With Claude 2025, a sample promoted workflow is to have AI implement a mock UI (mock.png), then give it Puppeteer to take screenshots, and ask it to reiterate till pixel-perfect. Once I get into the gist of it, it is actually less delegating and more system design, building a platform to automate the boring work away.

AI-first projects

I find that sometimes we give AI the most gnarly bits of the system to work on, bits that we don’t even want to look at ourselves, get disappointed at the mess, and declare AI is not ready yet. The last part is correct, they are junior, the eldest is around 3 in human years. Think about what most 3-year-olds do: they put things in their mouth, draw on walls, and occasionally produce something brilliant by pure accident. Sound familiar?

Greenfield projects where the code base is still small, and AI-friendly design principles (such as modularity, clear APIs, good documentation) are easy to enforce play on the strength of AI and negate some of its most significant drawbacks.

Some projects take it a step further. From their inception, certain parts of the software were destined to be written by AI.

The human code orchestrates the service's overall behavior and is kept segregated from the AI code. Think of the strategy or decorator design pattern on steroids.

The AI code is heavily modularized, has a lower standard of code quality, and can be rewritten (or rather, regenerated) at a moment's notice.

Requirements and tests are emphasized because they are the primary deliverable, not the expendable code. Their creation, in turn, can also be AI-assisted as part of the AI crew concept.

We do a fair share of custom HTML email templates for our customers. Crafting HTML and CSS is not the most thought-provoking activity, and the constant back-and-forth checks to support pixel-perfect design across browsers and email clients are nerve-wrecking. Needless to say, turning that module into an AI-first project was enthusiastically supported.

Brownfield projects

Yet not every project can be an AI-first project. Despite constantly adding new services and breaking down old ones, I reckon many of the critical code blocks reside in a handful of projects that have been around since the beginning of time and are lovingly referred to as legacy code.

These projects are challenging even to the best human developers. A single feature might span dozens of files, each with its own decade-old conventions. Every developer who touched it left their mark. Now it's a Frankenstein’s monster of coding styles. And AI output is a function of its input.

We can adopt an incremental approach ("Strangler Fig" Pattern):
  • Before code modification, give the system’s documentation an overhaul. AI can generate initial drafts of documentation and summaries. The effort can be complemented with interviewing long-tenured developers to capture "tribal knowledge" - AI transcripting is particularly helpful. Finish up with a manual validation and refine. A happy side effect is that this logic documentation is also fundamental in building an AI data analysis experience.
  • If the code base is a behemoth, the context window will be a problem, especially with my favorite Claude. Divide the “memory” into tiers. The most basic level is the immediate prompt and the ongoing conversation, which forms the memory of the task at hand. Second is the project memory - claude.md, .cursor/rules, and whatnot - this hosts key decisions, patterns, and learned codebase knowledge. This will be added to the context window of every conversation, so be really selective about what is put in there. The last layer is an overview of the project, but sometimes the entire code base because of intersystem dependencies. This is the cutting edge of automatic code generation at the moment, companies are looking into RAG and GraphRAG solutions to bridge this gap.
  • Once the ground knowledge is there, avoid a "big bang" rewrite. Instead, use AI to help understand and build interfaces and unit tests around specific modules of the legacy system. Gradually replace or refactor these modules, with AI assisting in understanding the old logic and integrating new components.

Conclusion

From 10-15% productivity gains in Level 1 to the promises and pitfalls of Level 4 autonomous agents, one thing is clear: AI is reshaping software development at every level. The path forward isn't about choosing between human creativity and AI efficiency. It's about finding the right blend for your context. There is a whole spectrum to choose from, starting with delegating boilerplate to AI while keeping complex problem-solving for yourself, to architecting AI-first systems, to teaching AI to understand your decade-old codebase. In the end, the developers who thrive will be those who build better systems, be it a code production engine or features. It has always been that way.

Thursday, May 29, 2025

LLM Agents Comparison

 I was designing an AI-powered workflow for Product Management. The flow looks somewhat like this.


The details of the flow are largely not relevant to the topic of this post. The important gist is: it is an end-to-end flow that would be too big to fit in most modern LLMs’ context window. As the flow has to be broken down into sequential phases, it lends nicely to a multi-agent setup with intermediate artifacts to assist communication between one agent to another. It was the AI as Judge part that reminded me, just as a developer cannot reliably test their work, an agent should not be trusted to judge its own work.

"Act as a critical Product Requirement Adjudicator and evaluate the following product requirement for clarity, completeness, user value, and potential risks. [Paste your product requirement here]. Provide specific, actionable feedback to help me improve its robustness and effectiveness."

(A) I have a multi-agent setup. (B) I have a legitimate reason to introduce more than one LLM to reduce model blind spots. The next logical step is to decide which model makes the best agent.

An empirical comparison

While composing the aforementioned workflow, I fed the same prompts to Claude 3.7, Gemini 2.5 Pro (preview), and ChatGPT-o3. All in their native environments (desktop app for Claude and ChatGPT, web page for Gemini). I first asked the agent to passively consume my note while waiting for my cue to participate. For each note, I wrote one use case where I think GenAI is beneficial to the productivity of a product manager. When the source of the idea was from the Internet, I provided a link(s). Once my notes were exhausted, I asked the agent to perform a search to find other novel uses of GenAI in product management that I might have missed. Finally, with the combined inputs from me and web search, I asked the agent to compose a comprehensive report. I ended up with three conversations and three reports. Below are my observations on the pros and cons of each. An asterisk (*) marks the clear winner.

LLM as a thought partner

This is probably the most subjective section of this subjective comparison. Feel free to let me know if you disagree.

ChatGPT:

I didn’t like ChatGPT as a partner. Its feedback sometimes focused on form over substance - whether or not I follow the feedback makes no difference to my work. ChatGPT was also stubborn and prone to hallucination, a dangerous combination. It kept convincing me that Substack supports some form of table, first suggesting a markdown format, then insisting on a non-existent HTML mode.

Claude *:

I dig Claude (with certain caveats explored below). Its feedback had a good hit-to-miss ratio. The future-proof section at the end of this post was its suggestion (the writing mine). There seemed to be fewer hallucinations too. Empirically speaking, I had fewer “this can’t be right” moments.

Gemini:

Gemini’s feedback was not as vain as ChatGPT’s but limited in creativity, and had I followed, I doubt it would have made my work drastically different. Gemini gave me a glimpse into what a good LLM can do, yet left me desiring more.

Web search

ChatGPT *:

The web search was genuinely good. It started with my initial direction and covered a good ground of internet knowledge. The output contributes new insights that didn’t exist in my initial notes while managing to stay on course with the topic.

Claude:

The web search seemed rather limited. Instead of including novel uses, it reinforced earlier discussions with new data points. More of the same. Good for validation but bad for new insights. Furthermore, Claude’s internal knowledge is limited to Oct 2024, it could list half a dozen acronyms of MCP, except Model Context Protocol!

Gemini *:

Gemini web search was good, it performed on par with ChatGPT, even faster.

Deep thinking

Aka extended thinking. Aka deep research. You know it when you see one.

ChatGPT *:

Before entering the deep thinking mode, it would ask clarification questions. These questions have proven to be very useful in steering the agent in the right direction and saving time. The agent would show the thinking process while it is happening, but hide it once done. Fun fact, ChatGPT used to hide this thinking process completely until DeepSeek exposed it, “nudging” the OG reasoning model to follow suit.

Claude:

Extended thinking was still in preview. It didn’t ask for clarification, instead uncharacteristically (to Claude) jumped head-on into the thinking mode. It also scanned through an excessive number of sources. Compared to the other 2 agents, it didn’t seem to know when to stop. In that process, it consumed an excessive number of tokens. This however is every much characteristically Claude. I am on the Team license, and my daily quota limit was consistently depleted after 3-4 questions. Claude did show and keep displaying the thinking process.

Gemini:

Deep research did not ask clarification questions, it showed the research plan and asked if I would like to modify. Post confirmation, it launched the deep research mode. Arguably not as friendly as ChatGPT’s approach but good for power users. The thinking process of Gemini was the most details and displayed on the most friendly UI among the bunch. However it was very slow. Ironically, Gemini didn’t deal with disconnection very well, even though at that speed, deep research should be an async task.

Context window

ChatGPT: 128-200k

The agent implemented some sort of rolling context window where earlier chat is summarized in the agent’s memory. The token never ran out during my sessions, but this approach is well-known for collapsing context where the summary misses earlier details, and the agent starts to fabricate facts to fill in the gap.

Claude: 200k

Not bad. Yet reportedly the way Claude chunks its text is not as token-efficient as ChatGPT. When the context window exceeded, I had to start a new chat, effectively losing all my previous content unless I made intermediate artifacts beforehand to capture the data. I lived in the constant fear of running out of tokens, similar to using a desktop in electricity-deficient Vietnam during the 90s.

Gemini *: 1M!

In none of my sessions did I manage to exceed this limit. I don’t know if Gemini simply forgets earlier chat, summarizes it, or requires a new chat. Google (via Gemini web search) does not disclose the precise mechanism it chooses to handle extremely long chat sessions.

Human-in-the-loop collaboration

ChatGPT:

ChatGPT did not voluntarily show Canvas when I was on the app but would bring it up when requested. I could highlight any part of the document and ask the agent to improve or explain it. There were some simple styling options. It showed changes in Git diff style. ChatGPT chart was not part of an editable Canvas. Finally, it did not give me a list of all Canvases. They were just buried in the chat.

Claude *:

Claude’s is called Artifacts. Claude would make an artifact whenever it feels appropriate or is requested. Any part can be revised or explained via highlighting. Any change would create a version. I couldn’t edit the artifact directly though. Understandably, there were also no styling options. Not only could Artifact display code, it also executed JavaScript to generate charts, making Claude’s chart the most customizable in the bunch. To be clear, ChatGPT could show charts, it just did so right in the chat as a response.

Gemini:

Gemini’s Canvas copied ChatGPT’s homework, all the good bits and bad bits. I had to summon it, no Canvas list, no versions but I could undo. Chart rendering however was the same as Claude's - HTML was generated. Gemini added slightly more advanced styling options and a button to export to Google Docs. Perhaps it was just me, but I had several Canvas pages that went from being full of content to blank randomly, and there was nothing I could do to fix the glitch. It was madness.

Cross-app collaboration:

ChatGPT:

ChatGPT has some proprietary integrations with various desktop apps but the interoperability is nowhere as extensive as MCP. For example, where the MCP filesystem can edit any files within its scope of permission, ChatGPT can only “look” at a terminal screen. MCP support was announced in April 2025, this might be improved soon.

Claude *:

MCP! Need I say more?

Gemini:

It is a web page, it is a closed garden with only Google Workplace integrations (“only” for many, this might be the number one reason to stick with Gemini). It is also unclear to me if Google’s intended usage for Gemini the LLM is via NotebookLM, the Gemini webpage, or the Gemini widget in other products. MCP support was announced but being a web page, it will be limited to remote MCP servers.

Characteristics

ChatGPT:

A (over)confident colleague. It is designed to be an independent, competent entity with limited collaboration with the user. other than receiving input. Excellent at handling asynchronous tasks. Also known for excessive sycophancy (even before the rollback).

Claude:

A calm, collected intern. It is designed to be a thought partner working alongside the user. Artifacts and MCPs reinforce this stereotype. It works best with internal docs. Web search/extended thinking need much refinement.

Gemini:

An enabling analyst. Gemini’s research plan and co-edit canvas place a lot of power into the hands of its users. It requires an engaging session to deliver a mindblowing result. Also because if you turn away from it, the deep research session might die, and you have to start again.

Putting the “AI Crew” Together

With that comparison in mind, my today AI crew for product management looks like this.

  • Researcher:
    • Web-heavy discovery: ChatGPT
    • Huge bounded internal doc: Claude
  • Risk assessor: Claude (Claude works well with bounded context)
  • Data analyst:
    • Claude with a handful of MCPs
    • Gemini if the data is in spreadsheets
  • Backlog optimizer: Claude with Atlassian MCP
  • Draft writer:
    • ChatGPT for one-shot work and/or if I know little about the topic
    • Claude/Gemini if I’m gonna edit later
  • Editor/Reviewer: Claude Artifacts = Gemini Canvas
  • Publisher: Confluence MCP / Manual save

Some rules of thumb

  • Claude works well with a definite set of documentation. If discovery is needed, ChatGPT and Gemini are better.
  • Between ChatGPT and Gemini, the former is suitable for less technical users who are less likely to engage in a conversation with AI.
  • In some cases, Claude is selected simply because the use case calls for MCP, and it is the only consumer app supporting MCP today. Also because of this, for tasks all 3 LLMs perform equally well, I go with Claude for later interoperability.

Future proof

When I finished the rules of thumb, I had an unsettling feeling, this would not be the end of this. As the agents evolve, and they do - neck-breaking fast, this guide will soon become obsolete. I don’t plan to keep the guide forever correct, nor am I able to. But I can lay out the ground rules to keep reshuffling and reevaluating the agents as new improvements arrive.

  • Periodic review - What was Claude's weakness yesterday might be its strength tomorrow (looking at you, extended thinking). Meanwhile, a model's strength could become commoditized, diluting its specialization value. Set a quarterly review cadence to assess if the current AI assignments still make sense. Even better if the review can be done as soon as a release drops.
  • Maintain a repeatable test - Develop a standardized benchmark that reflects your actual work. Run identical prompts through different models periodically to compare outputs objectively. This preserves your most effective prompts while creating a consistent evaluation framework. Your human judgment of these comparative results becomes the compass that guides crew reorganization as models evolve. If the use case evolves, evolve the test as well.
  • Build platform-independent intermediate artifacts - As workflows travel between AI platforms, establish standardized formats for handoff artifacts (PRDs, market analyses, etc.). This reduces lock-in and makes crew substitution painless. A good intermediate artifact should work regardless of which model produced it or which model will consume it next. In my use case of product management, some of the artifacts are stored in Confluence.


 

Sunday, May 11, 2025

Multi vs Single-agent: Navigating the Realities of Agentic Systems

In March 2025, UC Berkeley released a paper that, IMO, has not been circulated enough. But first, let’s go back to 2024. Everyone was building their first agentic system. I did. You probably did too.

We defined the systems into different maturity levels

Level 1: Knowledge Retrieval - The system queries knowledge sources to fulfill its purpose, but does not perform workflows or tool execution.

Level 2: Workflow Routing - The system follows predefined LLM routing logic to run simple workflows, including tool execution and knowledge retrieval.

Level 3: Tool Selection Agent - The agent chooses and executes tools from an available portfolio according to specific instructions.

Level 4: Multi-Step Agent - The agent combines tools and multi-step workflows based on context and can act autonomously on its judgment.

Level 5: Multi-Agent Orchestrator - The agents independently determine workflows and flexibly invoke other agents to accomplish their purpose.

You might have seen different lists, yet I bet that no matter how many others you have seen, there is a common element: they all depict a multi-agent system as the highest level of sophistication. This approach promises better results through specialization and collaboration.

The elegant theory of multi-agent systems

The multi-agent collaboration model offers several theoretically compelling advantages over single-agent approaches:

Specialization and Expertise: Each agent can be tailored to a specific domain or subtask, leveraging unique strengths. One agent might excel at generating code while another specializes in testing or reviewing it.

Distributed Problem-Solving: Complex tasks can be broken into smaller, manageable pieces that agents tackle in parallel or sequence. By splitting a problem (e.g., travel planning into weather checking, hotel search, route optimization), the system can solve parts independently and more efficiently.

Built-in Error Correction: Multiple agents provide a form of cross-checking. If one agent produces a faulty output, a supervisor or peer agent might detect and correct it, improving reliability.

Scalability and Extensibility: As problem scope grows, it's often easier to add new specialized agents than to retrain or overly complicate a single agent.

The theory maps beautifully to how we understand human teams work: specialized individuals collaborating under coordination often outperform even the most talented generalist. It is every tech CEO’s wet dream: your agent talks to my agents and figures out what to do.

This model has shown remarkable success in some industrial-scale applications:

Then March 2025 landed with a thud!

The Berkeley reality check

A Berkeley-led team assessed the stage of current multi-agent system implementation. They ran seven popular multi-agent systems across over 200 tasks, and identified 14 unique failure modes organized into three categories:

The paper Why Do Multi-Agent LLM Systems Fail? showed that some state-of-the-art systems achieved only 33.33% correctness on seemingly straightforward tasks like implementing Tic-Tac-Toe or Chess games. For AI, this is the equivalent of getting an F.

The paper provides empirical evidence for what many of us have experienced: today's multi-agent implementation is hard. We can’t seem to fulfill the elegant theory outside a demo, with many failures stemming from coordination and communication issues rather than limitations in the underlying LLMs themselves. And if I didn’t make the table above myself, I would think it was a summary of my university assignments.

The disconnection between PR and reality

Looking back, it's easy to see how the industry became so enthusiastic about multi-agent approaches:

  1. Big-cloud benchmarks looked impressive. AWS reported that letting Bedrock agents cooperate raised "marked improvements" in internal task-success and accuracy metrics for complex workflows. I am sure AWS has no hidden agenda here. It is not like they are in a race with anyone, right? Right?
  2. Flagship research prototypes beat single LLMs on niche benchmarks. HuggingGPT and ChatDev each reported better aggregate scores than a single GPT-4 baseline on their chosen tasks. In the same way the show Are You Smarter Than A 5th Grader works. But to be frank, the same argument can be used for the Berkeley paper.
  3. Thought-leaders said the same. Andrew Ng's 2024 "Agentic Design Patterns" talks frame multi-agent collaboration as the pattern that "often beats one big agent" on hard problems.
  4. An analogy we intuitively get. Divide-and-conquer, role specialization, debate for error-catching — all map neatly onto LLM quirks (limited context, hallucinations, etc.). But just with humans, coordination overhead grows exponentially with agent count.
  5. Early adopters were vocal. Start-ups demoed agent-teams creating slide decks, marketing campaigns, and code bases - with little visible human help - which looked like higher autonomy. Till reality’s ugly little details turn this bed of roses into a can of worms.

The Berkeley paper exposed these challenges, but it also pointed toward potential solutions.

Enter A2A: Plumbing made easy

Google's Agent-to-Agent (A2A) protocol arrived in April 2025 with backing from more than 50 launch partners, including Salesforce, Atlassian, LangChain, and MongoDB. While the spec is still a draft, the participation signal is strong: the industry finally has a common transport, discovery, and task-lifecycle layer for autonomous agents.

A2A directly targets 03 of the 14 failure modes identified in the Berkeley audit:

  1. Role/Specification Issues - By standardizing the agent registry, A2A creates clear declarations of capabilities, skills, and endpoints. This addresses failures to follow task requirements and agent roles.
  2. Context Loss Problems - A2A's message chunking and streaming prevent critical information from being lost during handoffs.
  3. Communication Infrastructure - The HTTP + JSON-RPC/SSE protocol with a canonical task lifecycle provides consistent, reliable agent communication.

Several vendors have begun piloting A2A implementations, with promising but still preliminary results. However, it's important to note that quantitative proof remains scarce. None of these vendors has released side-by-side benchmark tables yet—only directional statements or blog interviews. Google's own launch blog shows a candidate-sourcing workflow where one agent hires multiple remote agents, but provides no timing or accuracy metrics. I won’t include names here because I believe we have established earlier that vendor tests can be unpublished, cherry-picked and may have included proprietary orchestration code.

In other words, A2A to agents is like what MCP to tool uses today. And just as supporting MCP doesn’t make the code of your tools better, A2A doesn’t make agents smarter.

What A2A Doesn't Solve

While something like A2A will fix the failures stemming from the lack of a protocol, it cannot fix issues that do not share the same root cause.

  • LLM-inherent limitations - Individual agent hallucinations, reasoning errors, and misunderstandings. Just LLM being LLM.
  • Verification gaps - The lack of formal verification procedures, cross-checking mechanisms, or voting systems. Without verification, you're essentially trusting an AI that thinks 2+2=5 to check the math of another AI that thinks 2+2=banana.
  • Orchestration intelligence - The supervisory logic for retry strategies, error recovery, and termination decisions. Just the other day, Claude was downloading my blog posts, hit a token limit, tried to repeat, continued to hit the token limit, and looped forever.

Those are 11 out of 14 failure modes. These areas require additional innovation beyond standardized communication protocols. Better verification agents, formal verification systems, or improved orchestration frameworks will be needed to address these challenges. I would love to go deeper into the innovation layers required to solve the “cognitive” failure of multi-agent models in another post.

Conclusion

Multi-agent systems are not the universal solution for today’s problem (yet). Many are here and working, but only at a scale where no single-agent approach can reach. These multi-agent systems deliver potential benefits, but at a substantial cost, one that can easily exceed an organization's capacity to invest or find people with the relevant expertise.

Rather than viewing multi-agent as the holy grail to be sought at all costs, we should approach it with careful consideration—an option to explore when simpler approaches fail. The claimed benefits rarely justify the implementation complexity for most everyday use cases, particularly when a well-designed single-agent system with appropriate guardrails can deliver acceptable results at a fraction of the engineering complexity.

Still, the conversation of multi-agent systems will always be around, in increased frequency, especially in light of the arrival of protocols like A2A. MCP is still the hype now, organizations are busy integrating MCP servers and writing their own before turning their attention to the next big thing. A2A could be that. Better agent alignment could also be that. Or it requires a new cognitive paradigm to improve agents’ smartness. Things will take another 6 months to unfold, which is the amount of time since the MCP announcement.

Either way, what’s clear is that it is a long way for agentic systems to hit their theoretical performance plateau.

Sunday, April 20, 2025

Finally, a break

Hi guys,

The news was out last Friday. I am going to leave Parcel Perform. 

For a sabbatical leave. Between May and October.

Sorry for the gasp. You come here for drama.

This has been planned since last year. I originally planned to leave soon after the 2024 BFCM, after the horizontal scalability of the system became a solved problem. I would like to say permanently, but I have learned that nothing ever is - the scalability, not me leaving for good. But then there were such and such issues with our data source (if you know, you know), and AI took the industry by storm. So I stayed. Eventually though, I knew I needed this break.

I have been on this job for almost 10 years. The first line of code was October 2015. I thought I would be done and move on in 5 years! I have been around longer than most furniture in the office and gone through 4 office changes. A decade is indeed a long time. It is a wet blanket that dampens any excitement and buzz that comes out of the work. Things get repetitive. Except for system incidents, I have lost count of how many ways things can combine to blow off. I praise every morning to wake up to no new alert.

When I was 23, I was fired from my job, and I was unemployed for 6 months. More like unemployable. It could have been longer had my savings not gone dry. I wrote, read, cooked, rode, swam, organized events, and lived a different life I didn't know I should. It was the time I needed to recover from depression. It was the best of my life. I want to experience that one more time.

Upon this news, I received some questions, the most common ones are below.

Why are you leaving now? Is something bad happening?

I am still the CTO of Parcel Perform, just on sabbatical leave. And on the contrary, I think this is a good window for me to take a break. The business is in its best shape since the painful layoff in 2022. We are positive about the future, we are actively expanding the team for the first time in 3 years. The Tech Hub, of which I am directly responsible, has demonstrated that in the face of unprecedented incidents, we are resilient, innovative, and get things done. With the multi-cluster architecture and other innovations, we won't face an existential scalability problem for a long time.

In the last 6 months, we have invested in incorporating GenAI into our product. I believe we have the right data, tech stack, and an experiment-focused process, though only time can tell. To be frank, all the fast-paced experiments we are doing, known internally as POC, reminded me of all the things I loved about working here in the early days. Ideas are discussed in a small circle, stretched on a whiteboard, implemented in less than a day, and repeated. It has been so fun recently that I got cold feet. Perhaps I shouldn't take this break yet. But I am not getting younger, I am getting married, and soon will start a family. I won't have time for myself in a long time. It has to be now.

What will you do during the break?

Oh wow I'm gonna play so much computer game, my brain rots. I have been a vivid fan of Age of Empires II since the time there was only one kid in my neighborhood with a computer good enough to play the game. I am an average player, slow even, so perhaps we are looking at more losses than wins. But hey, it builds character.

I will host board game nights here and there. It's another long-lasting hobby of mine, and a perfect social glue for my group of friends. While I am at that, I probably want to up my mocktail game too. My friends are largely in my age bracket, so for the same reasons above, my ultimate goal is to have more quality time.

As far as dopamine goes, that's it. I am not planning for retirement after all. Can't afford that yet.

To be frank, the pretext of this break is that I want to work on my sleep, which has been less than ideal for a long time. I couldn't figure out a single one thing that could improve my sleep so it is gonna be a total revamp. Distance from work. White bed sheet. Sleep hygiene. Gotta catch em all.

I am probably still awake more than 12 hours a day though. I will be reading as much as I can, fiction, non-fiction, and whatnot. Real life is crazy these days; the distinction is getting vague. There are some long Murakami novels I want to go through. I find that reading his works in one go, or at least with a minimal pause, offers the ideal immersive experience.

What I read, I write. I hope I can find an interesting topic to write every month. If you are keeping track, the last few days have been quite productive ;) I am starting my first subtrack AI Crossroads because that's how I feel these days: an important moment in my life, our lives, that I cannot afford to miss. I am excited and confused. I am sure somebody out there is feeling the same. 

And I will pick up Robotics as a new hobby. As GenAI gets "solved", its reach will not stop at the virtual world. Robotics seems to be the next logical frontier where a new generation of autonomous devices crops up and occupies our world. 

Writing about these things, I'm already pumped!

Who will replace your role?

The good thing about cooking up this plan from last year is that I have had plenty of time to arrange for my departure. The level of disruption should be minimal. People won't notice when I am gone and when I am back. Or so I hope.

There isn't a simple 1 to 1 replacement. Parcel Perform is not getting a new CTO, and there will still be a job for me when I am back. Or so I hope.

As a CTO, my work comes in 3 buckets: feature delivery, technical excellence, and AI research.

Feature delivery is where we have the most robust structure. Over the years, we have managed to find and grow a Head of Engineering and two Engineering Managers. The tribe-squad structure is stable. We are getting exponentially better at cross-squad collaboration as well as reorganization. There is a handful of external support ranging from QA, infrastructure, to data and insight to ensure the troops have the oophm they need to crush down deliveries.

Technical excellence means making sure Parcel Perform tech stack stays relevant for the next 5 years. This is an increasingly ambitious goal. Our tech stack is no longer a single web server. The team is growing. The world is moving faster. But we have 4 extremely bright Staff Engineers. They each have spent years at the organization, are widely regarded for the depth of their technical knowledge, and are definitely better than me on my best day in their field of expertise. We have spent the last couple of months aligning their goals with the needs of Parcel Perform. The alignment is at its strongest point ever since we adopted a goal-oriented performance review system.

Lastly, AI research is the preparation of the organization for the AI future, across technologies, processes, and strategic values. While I will continue the research in my own time, there is now a dedicated AI team that has been made the center of Parcel Perform's AI movement. Despite the humble beginning, the team is 2x their size in the coming months and won't let us "go gentle into that good night" that is our post-apocalyptic lives with the AI overlords.

I think we are in good hands.

What will you do when you come back?

Honest answer, I don't know.

Also honest answer, I don't think it gonna be the same as what I am doing today. Sure, some aspects gonna be the same. 5 months isn't that long. Neither is it short. The organization will continue to grow and evolve to meet the demand of the market, the gap I leave will be filled. When I am back, the organization will undoubtedly be different from what it is today. I will have to relearn how to operate effectively again. I will need to identify the overlap between my interests, my abilities, and the needs of the new Parcel Perform. 

Final honest answer, I am anxious for that future, and it is the best part.