AI Took the Coding, Not the Judgment

AI Took the Coding, Not the Judgment

Where AI-Native Delivery Starts

AI was supposed to make coding cheaper. It made human coding uncompetitive at any price. That is the part of software delivery AI has taken, and where AI-native delivery starts. The rest of it AI has changed but not taken over: deciding what to build, getting from a vague idea to a working system, knowing when something works. I’ve written about how product management is changing under AI.

For as long as most of us have worked in this industry, software delivery was measured in units of work: hours or story points. Companies bought coding from a vendor or staffed it in their own IT team. That is the market that just disappeared. A team built to supply coding hours has lost most of what it used to deliver. A team built around judgment, deciding, specifying, and reviewing has the same work it had before, just done faster. This post is about the difference between the two approaches, and why AI-native delivery is the second one.

The hard part for the first kind of team is that coding hours were never just a price. They were an ideology and a business model. For a vendor, they decided how it hired, how it created seniority, how it quoted, how it talked to clients, and why it rarely promised results. For an internal team, the same logic sits in headcount plans and budget rounds instead of invoices. The way the team hires, grades, quotes, and promises stays exactly as it was.

Coding hours were never just a price. They were a business model.

Two Kinds of Expertise Behind AI-Native Delivery

I think about expertise as two things, and people usually mean only the first one. One is craft: being senior in a discipline. I’ve spent nearly thirty years in sales, business development, and product management. Product is where I’ve spent the most time, including running it for a Nokia business unit doing over two billion dollars a year. At F-Secure my team built a Salesforce security product that Salesforce itself is only now, nearly a decade later, shipping natively.

Branden Crawford, our CTO, brings the same seniority to software development and architecture. Neither of us is an expert in our clients’ industries, finance, heavy machinery, or mining, and we don’t pretend to be. Working inside a finance company doesn’t make a developer fluent in how the trading desk runs, either. That is worth sitting with if you build software in-house: proximity to the business is not the same as knowing it. Inside a company, discovery is one call away from a colleague, and it is the call that rarely gets made.

Process Means Asking, and Asking Again

The second kind of expertise is process. It means sitting down with the people who have the knowledge, the traders, the plant managers, and turning what they know into something we can build against. Call it discovery. Much of what they know has never been written down, and some of it has never been said out loud. Not pretending to already understand the domain is what forces us to go ask. Good process work often gets further than the client’s own team has, and that’s not because we’re smarter. We aren’t carrying years of internal history; we are working to a budget and a date rather than a backlog, and we can’t skip discovery. An in-house team can do all of this. The backlog is what usually gets in the way: when there is always a next ticket, discovery is the step teams skip.

Not pretending to already understand the domain is what forces us to go ask.

The five-step process we run is built on that discovery. Process expertise is what makes AI-native delivery hold up in practice. The tooling is secondary.

In August, we exhibited at the iVT Expo in Chicago, an off-road heavy machinery fair, not a software one. Most AI on the floor showed what the technology does. We went the other way. Before building a demo, we asked partners, vendors, and operators four questions: What do you need to do your work better? How can we help you sell more? How can we help you reach better quality? Where are the savings in operations? The answer we kept hearing was diesel. The demo showed how a 50-machine fleet would save about $100,000 a year in fuel from data the machines already record, coaching operators toward the engine’s sweet spot and rewarding them for staying there. In a follow-up call, a contact from the fair introduced us to his team as the one AI conversation there that had made sense to him. That shouldn’t be rare.

AI-Native Delivery Is a Process, Not a Plugin

I should say what I mean by AI-native delivery. Most of what vendors sell under that name is a team that bought some licenses. The same goes for a lot of what companies announce internally. Ours wasn’t bought. When Ilpo Niva, our Chief AI Officer, built the first version of our AI tooling, it didn’t make anything faster. Branden took it into real internal and customer work, hit the walls, sometimes had to abandon the tooling mid-task and build by hand, and reported back. Ilpo changed the tooling. Branden tried again. The first rounds cost us time. Then it reached parity with manual work, and now it is ahead in most everyday development tasks. The tool stack we run today came out of that loop, and it is the reason I trust the definition that follows.

The test is simple. Take any piece of work on a project and ask whether it matters who picks it up, a person or an AI agent. On an AI-native team it doesn’t matter. Both work from the same issues, the same board, the same pull requests, and the same reviews. There is no separate AI track that somebody has to reconcile later. What does differ is who decides. AI drafts, and a person approves everything that lands, with silence never counting as a yes.

What Holds It Together

The part clients rarely see is what holds it together: every task traces back to a line in the brief the client signed. We run five steps on every engagement: brief, spec, plan, code, review. When I first wrote the steps up, the fifth was deployment; we have since moved review into the core and left deployment outside it, because an app store, an Azure tenant, and a client’s own cloud have almost nothing in common. We build nothing until the AI has read the brief and the plan the work belongs to. Test-first development, code review, and security review run inside that flow, whether a human or an agent wrote the code. That is the AI-native engineering underneath the delivery. Every engagement feeds what it learned back into the process. Nothing in those five steps needs a vendor. An internal team can run exactly this.

I think about it the way I think about security. If it wasn’t designed in, you can’t glue it on afterward. Bolt AI onto an old delivery model, run by a vendor, or in your own engineering department, and you get the same implementation, faster, if you are lucky. Build the process AI natively, and the implementation hours go away.

Bolt AI onto an old delivery model and you get the same implementation, faster, if you are lucky. Build the process AI natively, and the implementation hours go away.

Cost-First Buying Has a Limit

Buyers are used to pricing vendors by the day rate. Hiring managers price their own teams by the headcount grade. It is the same arithmetic, and I’ve watched it go wrong the same way on both sides. Two junior developers look cheaper than one senior, so people buy two or hire two. What they usually get is more coordination, more mistakes to catch, and a total bill close to or above what the senior would have cost, for less output. We sell on a time-and-materials basis, too, so this isn’t a complaint about day rates. What misleads is comparing rates without comparing what a day produces. The day rate is real. The savings mostly aren’t.

The day rate is real. The savings mostly aren’t.

Where the Day Rate Breaks

A day-rate model breaks down the moment the work stops being well specified. A team built around implementation hours has no slack to absorb ambiguity. When the requirements turn out to be incomplete, the gap has to go somewhere. With a vendor, it becomes a change order. In-house, it becomes a quarter-long delay that never shows up on an invoice, which is why nobody fixes it. McKinsey and the University of Oxford looked at more than 5,400 IT projects. The large ones ran 45 percent over budget and delivered 56 percent less value than planned, with unclear requirements as a main cause. That was 2012, before anyone could build the wrong thing overnight.

We registered on SAM.gov about three years ago, which lets a company bid on US federal contracts. Often, cost is the most important decision-making criterion. Before we bid, we talked to companies already working that market.

You win the initial bid on price, almost always. Value arguments don’t move a competitive bid much. If that were the whole story, the most aggressive bidders would lose money and wash out, and the market would settle into something normal.

It doesn’t, because of the kicker. Once you’re the awarded contractor, change orders aren’t competitively bid anymore. The winning strategy is bidding at zero margin, or close to it, and making the money on change orders afterward, priced however you want, because nobody bids against you. We chose not to play this game.

Once we understood that, we saw the same pattern across low-cost work generally, not just federal contracts.

There is nothing wrong with a day-rate shop; it just stops working once AI does the work the day rate was priced for. The vendor in that story is a different animal. Call it the change-order vendor. Bidding at zero margin to win, then making the money on uncompeted change orders, isn’t a pricing decision. It’s a business model, and it tells you what the vendor will protect when something has to give. Every vendor is in business for a margin, us included. The difference is what each one sacrifices first. A low-bid vendor never sacrifices its margin. A vendor that lives on its reputation never sacrifices its reputation. Customer satisfaction is optional for only one of them.

When something has to give, the low-bid vendor keeps its margin. The reputation vendor keeps its reputation.

How to Read a Low Bid, or a Low Estimate

I tell everyone in this industry to think about the end user first. This is the rare exception. When a bid looks too good, think about the vendor.

Four questions about the vendor, not the product

  • What’s in this for the vendor? Every company needs a margin. If you can’t see theirs, it’s coming from somewhere you haven’t looked.
  • Why so cheap? Bench employees they don’t want to let go is a common answer. It’s also temporary. Once the oversupply is over, the appetite to deliver under the market price is gone.
  • What happens when a better-paying customer shows up? This only matters if you couldn’t see their margin in the first question. A vendor earning a fair margin on every account has nothing to rebalance. A vendor who underbid you does, and you’re the account it rebalances against.
  • Is cheap the business model? Some vendors swap in a cheaper team once you’ve committed. Others start overcharging on change orders. Once you’re hooked, fighting back is hard.

Don’t ask these on a sales call. Read the proposal and the contract. What are they obliged to do when the scope moves? Is there anything like a satisfaction guarantee that shows their intent? None of this argues for paying more. It argues for a price where both sides make money, because that is what keeps the team you were sold on the project. Maybe the team that impressed you in the sales meetings had quietly become a different team by month four. Maybe someone asked you to sign off on work you knew was below the agreed-upon threshold. Either way, you have already seen what that vendor sacrifices when something has to give. It wasn’t the margin.

Ask yourself what the vendor will sacrifice if it has to: its margin, or your satisfaction.

Build or Buy, After AI

The same four questions work on an internal estimate, and they are harder to ask because nobody is selling. A pattern I have seen more than once: a business owner can’t get their IT department’s attention, brings the work to us, and we quote it. Then IT picks it up, and we hear that IT will build the project in-house after all.

Sometimes the reason is price, and sometimes that is legitimate. IT doesn’t bill every internal cost back to the business, and the company may already own the licenses. Sometimes capacity frees up. Or the price never comes up. Sometimes the IT team would rather own the work than have an outsider question its assumptions, which is a human reason, not a bad one. But it is a reason the business owner should hear out loud. The test is the same as for a vendor: if you can’t see how the internal team does this more effectively, something else is motivating the number.

What do you know about the team that gave the work estimate? Is the estimate low due to confidence, or because it is the only number that gets the project approved? What happens when a higher-priority project lands on the same people, because one will? An internal team that underestimates to get funded is running the change-order vendor’s playbook with the invoice removed. The project pays anyway, in time instead of money.

When In-House Is the Right Answer

So, when is in-house the right answer? Before AI, I would have said: whenever the software is the company’s core, or has to run and evolve for years after launch. Or when owning the data and the people matters more than speed. The mirror image holds too: if the company has no steady demand for the skills, cannot attract and keep them, or needs an outside view, a vendor is the better answer. That logic still holds, but it now answers only half the question.

There are two questions, not one: what should you build in-house, and what can you. The second is new. The answer depends on how far the internal team has come with AI-native delivery. No team, ours included, is the most reliable judge of its own capability. Willing is not the same as capable, and a team that hasn’t worked this way yet may not know what capable looks like. I see it in clients all the time: people so busy running alongside the bike that they never have time to jump on it and ride. The business owner should decide what the company builds in-house. Whether the team can build it is a question of evidence: ask for a recent project delivered this way, not a plan to deliver one.

Willing is not the same as capable.

The Cost Premium Is Shrinking

To be fair to the in-house option, an outside vendor has always carried a cost premium. Somebody pays for the cost of sales, and you end up funding two human organizations, the vendor’s team on top of your own. What you got for it was flexibility in resourcing and an outsider’s view. My claim is that the cost premium is narrowing fast, because most of it was the overhead of the vendor’s human team. Open a job, interview, hire, buy the laptop and the licenses, train the person, and wait for them to ramp up on their own systems and yours. The first invoice for their work might arrive four months after the job was posted.

While they work, someone runs payroll, insurance, expenses, and reviews for them, and once the team is big enough, someone manages every eight of them. None of that is billable, so all of it lived in the vendor’s margin, which is to say in your price.

AI capacity carries almost none of that overhead. You can add capacity in an afternoon, and adding more of it doesn’t mean adding another manager. So, the part of the premium that was a second human organization is the part that shrinks. What remains is the open question: whether the vendor gets more out of AI than your own team does. In our experience, we often do, though not always. A team built AI-native from the start gets more done than a team that added AI to a process designed for people. On many projects, that difference more than covers the remaining cost premium.

Why a Guarantee Needs a Reputation Vendor

Once AI does the well-specified coding, what a buyer still needs from a vendor is someone who will stand behind the result when the work is ambiguous. That is where two kinds of vendors lose, and they lose differently. The change-order vendor can’t stand behind anything, because its margin depends on the scope moving. The honest day-rate shop wants to, but the volume of its business ran on is gone. Neither can offer a guarantee, and the reason is not that they haven’t thought of it.

Change orders themselves are normal business. Scope moves, and when it does, the vendor prices the extra work and the client pays for it, with us as much as anyone. The change-order vendor’s trick is the one from the SAM.gov story: the first bid is priced to the bone, and nothing after the award faces competition. The day rate may be fixed, but whether a change takes ten days or thirty is the vendor’s number. Customers argue with it, and often they are right, but they have already chosen the vendor, and there is not much they can do. The margin lives in that number, and so does the tension you feel on those projects.

A satisfaction guarantee is a different approach, and it’s what we follow. Ours is a promise about our own work: if a week isn’t up to the client’s expectations, we fix it, or they don’t pay for it. That is a trust-based promise. It only works between two parties who expect good faith from each other every week. A project running on an unhealthy pricing model or outright deceit runs on tension instead. You cannot lay a trust-based guarantee on top of tension. That is why you rarely see a satisfaction guarantee from a low-bid vendor, and why it takes a reputation vendor to offer one.

The Day-Rate Shop’s Squeeze

The honest day-rate shop has a different problem. Its business was volume: enough implementation hours flowing through to make the day rate work. AI took most of that volume, and two things now squeeze the day-rate shop at once. The well-specified work it used to sell has been taken over by AI. The ambiguous work left over is exactly what a team built on low-cost, high-volume implementation hours can’t deliver without change orders. That would turn it into the change-order vendor it never wanted to be.

Some day-rate shops are restructuring around AI-native delivery, and it means rebuilding pricing, staffing, and sales incentives at once, which is why most haven’t. In-house, it is harder still: nobody ever asked the internal team for a guarantee, so there is no contract to renegotiate, only a habit to break.

An Unpopular Take

The change-order vendor is a product of how buyers buy. When procurement scores cost above everything else and has no adequate way to score judgment, it selects for the bidder willing to go to zero. Then it blames the project for the change orders that the selection guaranteed. Budget processes do the same thing to internal teams when they fund headcount by grade and cannot fund judgment. If you want a different kind of team, change what you reward. But if you can’t explain how the people building your software make their money, or make their numbers, don’t sign.

There is a personal version of this split too, what it means for the people whose careers started in the coding AI just took. I’ll write about that separately.

If you can’t explain how your vendor makes money, don’t sign.

We Put Our Own Money on It

A guarantee only works if the underlying odds are good. That’s the part a low-bid vendor can’t borrow. Ours, described below, works because AI-native delivery rests on the craft and the discovery, and those keep the odds good in the first place.

We break engagements into weekly deliverables now, not the two-week sprints most of the industry still runs on, because an AI-native delivery team moves faster than that. If a client is dissatisfied with a given week’s work, they tell us in writing within seven days. We meet, look at the specific concern, and get a fair chance to fix it. If, after that honest attempt, they’re still dissatisfied, we refund the full fee for that week.

We have never had to pay it. What it has done, three times that I can think of, is get a client to raise something they were hesitant about, a week in rather than a quarter in. Most of those turned out to be expectation gaps rather than defects. Every one was cheaper to close early than it would have been after the gap widened. The refund is the backstop. What the guarantee really buys is a client who says something in week one instead of week six.

The refund is the backstop. What the guarantee really buys is a client who says something in week one instead of week six.

You don’t need a contract to run this. A team lead can ask the people they build for, every week, whether the work was what they wanted, and mean it. The money is what makes it credible from a vendor. Inside a company, asking and acting on the answer is where most of the value lies.

Nobody can promise to always get it right. What we can promise is that if we fail, the person paying doesn’t carry the loss alone. That’s our 100% satisfaction guarantee: if you don’t like it, you don’t pay.

If you want to talk about how this would work on your project, reach out. And if you run an in-house team and think I have this wrong, I would like to hear it.

  • Mikko, Co-Founder and COO of A-CX has a background in driving innovation and building award-winning products and services. With extensive experience at Nokia, Microsoft, and F-Secure, Mikko has leveraged technology to create impactful solutions. Mikko’s career exemplifies a deep understanding of business dynamics and a passion for driving growth.

    COO, Co-Founder