Point of view

Why AI projects fail, and what a service business should do first

Most AI projects fail for a reason that has nothing to do with AI. Here is the pattern, what the research says, how it shows up in an HVAC shop, and the order I would do things in instead.

Every week another vendor calls an HVAC owner with an AI product. A voice agent that answers the phone. A chatbot for the website. A tool that writes follow-up texts. Some of these are good products. Most of the projects built on them still fail to pay back.

That isn’t because the software is bad. It is because the software gets chosen before anyone has looked closely at the work it is supposed to do.

The pattern

It goes like this. The owner feels a problem: phones going unanswered, estimates going cold, the office drowning in paperwork. A vendor offers a tool that sounds like it solves that problem. The tool gets bought and switched on.

Three months later the numbers look about the same. The voice agent books jobs, but some go into the wrong customer record. The follow-up texts go out, but half the open estimates were already sold or dead, so customers get chased for jobs that are done. Nobody measured where things stood before, so nobody can say whether anything improved.

The tool did what it was built to do. It was pointed at the wrong problem, or at the right problem with the wrong data underneath.

What the research says

This isn’t just my opinion. In 2024, RAND published a study based on interviews with 65 experienced data scientists and engineers about why AI projects fail. The report opens with an estimate that more than 80 percent of AI projects fail, about twice the rate of IT projects without AI. Read the RAND report.

RAND’s interviewees named five root causes. The most common was the first one on this list:

  • The people running the project misunderstood or miscommunicated what problem it was meant to solve.
  • The organization didn’t have the data needed to train or run the system properly.
  • The team focused on using the latest technology rather than on solving a real problem for the people who would use it.
  • The organization lacked the infrastructure to manage its data and put the system into use.
  • The technology was applied to a problem too hard for it to solve.

In 2025, a team at MIT published a report on generative AI in business, based on a review of over 300 public AI initiatives, 52 interviews and 153 survey responses. Its headline finding was that 95 percent of the organizations it studied were getting zero measurable return from their generative AI spending. Read the MIT report.

That MIT figure deserves some care. The report is preliminary, the sample is mostly large companies, and critics have questioned how “return” was measured. I wouldn’t lean on the exact number. But the finding points the same way as RAND’s: the gap was not about model quality. The projects that paid off were the ones fitted to how the business actually works.

Neither study looked at HVAC contractors. Both describe what I see when I look at how an HVAC shop runs.

How it shows up in an HVAC shop

The missed-call problem that was really a lunch problem

An owner knows calls are being missed and buys an after-hours answering tool. When you pull the call log, a large share of missed calls happen between 11:30 and 1:30, when two of three CSRs are at lunch together. After-hours coverage helps, but the biggest leak was a staffing schedule. That fix costs nothing.

The voice agent that couldn’t find the customer

A voice agent books a job by looking the caller up in ServiceTitan. If the same customer exists three times, with a landline on one record and a cell on another, the agent creates a fourth. The member discount doesn’t apply, the tech arrives without the equipment history, and the office cleans it up by hand. The tool is fine. The records underneath weren’t ready for it.

The follow-up that chased sold jobs

Automated estimate follow-up is one of the most valuable things a shop can add. But if comfort advisors don’t mark estimates as sold or dismissed, the sequence texts customers who already bought, or bought from someone else. It annoys customers and teaches the team to ignore the tool.

The dispatcher nobody wrote down

Many shops run on one dispatcher who knows which tech can handle which equipment and which customer is difficult. No software can automate what was never written down. The first step is a spreadsheet, not AI.

What to do instead

The order matters more than the tool. This is the order I work in.

  1. Map the work as it actually runs. Talk to the people who answer the phones and run the calls. Follow real jobs from first ring to paid invoice. Write down every handoff and every workaround.
  2. Measure the leaks. Put a dollar figure on each one, with the arithmetic shown, so you can check it. Missed calls a week, times the share that are real new jobs, times close rate, times average ticket, times 52.
  3. Fix the handoff, sometimes without AI. A lunch rota, a rule for marking estimates, a cleanup of duplicate records. Some of the best fixes are free.
  4. Then choose the tool. Now you know exactly what it has to do, what data it will read and how you will know it worked.
  5. Measure against the baseline. Compare the same numbers before and after. If they don’t move, stop paying for it.

Where AI does earn its keep

I’m not against AI. I build with it. In a service business it pays off when the job is repetitive, the rules are clear and the data is clean. For example:

  • Answering overflow and after-hours calls and booking straight into the schedule, once customer records are clean enough to find the caller.
  • Turning a tech’s voice notes into a clean job summary and invoice lines, once the office has agreed what a complete job record looks like.
  • Drafting estimate follow-ups that mention the actual equipment and options quoted, once estimate statuses are kept up to date.

Each of those comes with a “once”. That is the whole argument.

Five questions to ask before you buy

If a vendor is already on the phone with you, these questions separate a tool that fits from one that will sit unused.

  1. Which of my numbers will this move, and by how much? A good vendor names the number: missed calls, booked-call rate, close rate on estimates. A vague answer about efficiency is a warning.
  2. What does it need from my data? Ask which fields it reads in ServiceTitan or your CRM, and what happens when a customer exists twice.
  3. What happens when it gets it wrong? Every tool will misbook a job or misread a caller. You want to know who notices, and how fast.
  4. How will we know in 90 days whether it worked? If nobody measured the starting point, nobody can answer this.
  5. Who owns the setup if we stop paying? Accounts, call flows and scripts should be in your name.

None of these are hard questions. They are just rarely asked before the contract is signed.

If you want to know where your shop stands

This is what my diagnostic, Heimdall, does: two weeks tracing how the work really flows, a ranked list of leaks with the arithmetic shown, and a plain answer on what to fix first, including what not to automate.

You can read a full sample report, or book a 20-minute call and I’ll tell you whether a diagnostic makes sense for your shop.

Sources

Next step

Want to know where your shop stands?

  • You tell me what is going on in the shop.
  • I tell you whether a diagnostic makes sense.
  • No pitch deck. If it is not a fit, I will say so.

Email: [email protected]
Or send a message · Monday to Friday, 8am to 6pm Eastern

Pick a time for a 20-minute call. The calendar shows your local time.

Open the booking calendar