This week's news: Seventy four percent of customer facing agents already pulled back out, ninety five percent of organizations getting nothing measurable back, ten courses teaching people how to sell you a build, a hiring bar still screening for tools, and a wire that clears before lunch on a voice that never existed.
Read the whole thing in three minutes, or go deeper where it touches your business. We talk about the market here, not about ourselves.
Every article comes with a working tool. It runs right here in the page, it is free, and we do not need your information to use it. No signup, no download, nothing to fill in first. That holds for every article in every issue, back issues included.
Send it when it hits. Every article here has its own link; the send button next to each one copies it, ready to text or email to the person who needs it. Back issues stay in the archive and their tools keep working, so nothing you forward goes stale.
Want a reminder when the next one is up? Leave an email and we send one note a week. That is the whole program. We will not flood your inbox.
Sinch surveyed 2,527 senior decision makers across ten countries in January and February and published in May. Of the enterprises that got an AI customer communications agent into production, 74 percent have already rolled it back or shut it down. Among those describing their guardrails as fully mature the rate is 81 percent. 62 percent have agents live and 98 percent are still increasing spend. Gartner separately predicts 50 percent of companies that credited AI for headcount cuts will rehire those functions by 2027 under different titles.
The cost is not the license. It is the cut you made in month one. Severance, a rehire twelve months later at whatever the market charges by then, the recruiting, the ramp, and every customer who spent the gap talking to something that could not help them and could not say so. The saving was booked the day you cut. The bill arrives as a rollback, in a different quarter, which is exactly why nobody in the room connects the two.
Do not answer a rollback with a bigger model. The ones getting pulled are not failing at language, they are failing at governance. Do not cut headcount until the thing has run a full quarter in production without an incident. Do not read a low rollback rate as good news, because the number goes up at the companies with the best guardrails. They are not failing less. They are catching failures sooner.
Read the 81 percent again. The rollback rate is higher at the companies with the most mature controls, which looks like a paradox and is not one. Those companies are not failing more often. They are seeing the failure while it is still a rollback rather than a lawsuit. The organizations with weak controls are running agents that have already said something they should not have, and nobody has read the transcript yet. A low rollback rate is not evidence that yours is working. It is evidence that nobody is checking.
Before anything goes live, write three things the agent may never do without a person: move money, make a commitment, or take any action you cannot reverse. Then write the trigger that hands the conversation to a human, name the person who reads its transcripts every week, and hold headcount flat until it has run one full quarter with no rollback. The trade-off is that you carry the cost twice for a quarter. Against a severance and a rehire at market, one quarter of double running is the cheap version.
Sinch, The AI Production Paradox: a 2,527 person enterprise survey on why live AI agents get pulled back out · Sinch: the chapter carrying the production base rate and the fully mature guardrail number · Gartner: the February prediction, from 321 service leaders, that half of AI attributed cuts get rehired by 2027 · UC Today: independent reporting on the rollback finding
MIT's Project NANDA published The GenAI Divide in July 2025, built on 52 structured interviews, 153 survey responses from senior leaders and a review of more than 300 publicly disclosed initiatives. Against 30 to 40 billion dollars of enterprise investment it found 95 percent of organizations getting zero return and 5 percent extracting real value. Of organizations evaluating enterprise systems, 60 percent evaluated, 20 percent reached a pilot and 5 percent reached production. Tools bought from specialist vendors succeeded about 67 percent of the time. Tools built internally succeeded about a third as often.
The cost is not the pilot budget. It is two quarters of calendar, the integration hours your own people burned instead of doing their jobs, and the internal credibility the next attempt now has to buy back at a premium. Pilots get designed to impress the room rather than to move a number, so they end the only way they can, in a deck about learnings. The number was never defined, which means the pilot could not fail and could not succeed.
Do not run a second pilot on a newer model. That is the same clock restarted with the same missing scoreboard. Do not accept a demo as evidence, because the demo is the part they have actually finished. Do not build it internally because it looks cheaper, when the research puts internally built tools at roughly a third of the success rate of bought ones and your people are not sitting idle waiting for the work.
Handle this figure carefully, because it is famous and thinner than its fame. It is a version 0.1 preprint, the sample is 52 interviews and 153 responses, the unit is organizations rather than pilots, and it has been challenged in print. None of that rescues the pilot sitting in your business, because the mechanism it describes takes about four minutes to check yourself. Ask what number the pilot was supposed to move and by when. If nobody can answer, you are inside the 95 percent whatever the sample size was.
Before anything starts, write the one number the pilot has to move and the date it has to move by, and put both in the same document as the invoice. Then ask the vendor for a customer who has been in production for twelve months, and call that customer yourself. A vendor with real deployments hands you a name by the end of the day. A vendor selling demos schedules another demo. The reference call takes twenty minutes and it is the only part of diligence nobody can rehearse on your behalf.
MIT Project NANDA, The GenAI Divide: the July 2025 preprint behind the 95 percent figure, and its actual method · Virtualization Review: independent reporting carrying the correct sample sizes · Fortune: the coverage that put the number into general circulation, and the narrower claim it made
A ranked roundup of paid programs teaching people to start an AI automation agency was updated on September 6. It lists ten, each with a named operator behind it. The page closes with a section on why its author prefers a different business model entirely, and carries no affiliate disclosure. Search engines still index the same address under an older title claiming 34. There is no independent, editorially governed ranking of these programs anywhere, and no credible measure of the market's size, because the category is too new and too informal to have been measured.
The cost is not the course. It is who shows up in your office. The person pitching you a build may be six weeks out of a program that taught them to put a retainer on somebody else's tool, and the pitch sounds good because the pitch is what the course actually sold. What you get is a build that works in the demo, breaks on your real data, and cannot be repaired by the person who sold it, because they never wrote it.
Do not ask whether they have done this before. That was module three and the answer was written for them. Do not accept a case study in place of a client you can telephone. Do not let price be the test, because a build nobody can maintain is expensive at any number, and the cheap quote is usually the one with the shortest distance between the course and your invoice.
The tell is the shape of the answer, not its content. Ask what the system does when the data is messy. Somebody who built it answers in a sentence, because they have watched it happen and it ruined a week. Somebody who bought the answer sends a deck. Ask them to name the underlying tool their work sits on, and what happens to you if that tool changes its pricing or its interface. A builder names it immediately and has already thought about the second half. A reseller treats the question as rude.
Three questions and one phone call, in that order. What does it do when the data is messy. What is it built on, and what happens to me if that changes. Who is the last client still running it, and may I call them today. Then make the call yourself rather than having it arranged for you. The trade-off is that this adds a week, and good vendors will not mind, which is itself the test. Anybody who treats a reference call as an obstacle has already told you what the reference would say.
Ippei.com: a ten entry ranked roundup of paid AI agency programs, by an operator who sells a competing model at the foot of the page · AI Profit Boardroom: another roundup in the same genre, with no disclosed method · The Pivot Wave: a third, ranking many of the same programs in a different order
PwC's 2026 Global AI Jobs Barometer, published June 15, analyzed more than a billion job advertisements across 27 countries. Roles being professionalised by AI are growing twice as fast as roles being democratised by it, with 42 percent faster wage growth since 2021. New tasks added to AI exposed roles are 2.5 times more likely to depend on judgment, empathy and creativity. The AI skills wage premium reached 61.9 percent, up from 57. Entry level roles most exposed are seven times more likely to demand senior skills, and grew 35 percent since 2019 while other entry level roles fell 10.
The cost is not the salary. It is month five. You automated the documented, repeatable work, so what stays with a person is the exception, the judgment call and the conversation nobody wants to have. Then you hire against a bar you never updated, screening for tool proficiency and output volume, and the role fails on the parts you never tested for. You do not find out at the offer. You find out after the ramp, and you pay for the search twice.
Do not add an AI skills line to the job post. It selects for people who can name tools over people who can decide. Do not interview from the resume when the resume describes work the software now does. Do not price the offer against the tool list, because the tools are a week of training and the judgment is the entire hire.
The market split and your posting did not. PwC is reading a billion advertisements rather than asking anybody their opinion, and what those advertisements show is two tracks moving apart: work being professionalised by AI, where wages climb faster, and work being democratised by it, where they do not. The same split runs through your own payroll and it happened without a decision. Every task you automated moved a person up a track. Nobody rewrote the job description, so the hiring bar is still measuring the half you gave away.
Take your two hardest open roles and write down the three decisions that person will actually own. Not responsibilities. Decisions: the calls they make alone, with money or a customer on the other side. Build the interview around those three, ask for one they got wrong and what it cost, and price the offer against the judgment rather than the tool list. The trade-off is an afternoon per role and a posting that is harder to write. It also makes the fifth month survivable, which the posting currently does not.
PwC 2026 Global AI Jobs Barometer: a billion job advertisements across 27 countries, and the two tracks it found · PwC newsroom: the headline findings and the wage premium in one page · Resume Templates: 1,005 hiring managers on the skills they screen for, and where collaboration actually ranks · HR Dive: independent reporting on the same hiring manager survey
A survey of 1,533 corporate finance professionals in the United States and United Kingdom found 53 percent had been targeted by a deepfake financial scam and 43 percent fell victim. 87 percent said they would make a payment if called by their chief executive or finance chief. 57 percent can execute a transaction with no second approval. A separate five country survey put the average loss at 450,000 dollars, above 603,000 in financial services. The FBI logged 24,768 business email compromise complaints in 2025 at just over three billion dollars, and more than 22,000 complaints citing AI at 893 million.
The cost is not the wire. It is the control you never wrote, and the quarter that follows. A familiar voice has always worked as identity inside your business, so an urgent request in that voice skips every control you own, because your controls were written for documents and logins. The money leaves the same day and your bank cannot recall it after the cutoff. Then you have a controller who spends three months second guessing every legitimate request, which costs you in a currency nobody invoices.
Do not train people to hear the fake. That stopped being a human skill about a year ago, and telling your staff otherwise moves the liability onto the person least able to carry it. Do not verify on a number supplied during the call. Do not let the rule apply to everyone except you, because you are the voice being cloned, and the exception is the entire attack.
Put the two numbers side by side. 87 percent would pay on a call from the boss, and 57 percent can send the money with nobody else's approval. That is not a technology problem. It is an authority problem with a microphone pointed at it. The attack does not have to beat your systems. It has to reach one person who is allowed to act alone and who has been trained their entire career to treat urgency from leadership as the thing you do not question. Every control you own was written for a document.
Write one rule and date it. Any request to move money, change bank details or send payroll data gets verified by calling back on a number already in your records, never a number offered during the call. Add a second approval above a threshold you set yourself. Say in writing that nobody is penalized for making you wait ten minutes, and mean it the first time it happens to you. The trade-off is ten minutes on a genuine payment. Against a same day wire your bank cannot recall, that is not a trade-off.
FBI IC3 2025 Annual Report: the federal tally, including business email compromise and the first sizable count of AI enabled fraud · CFO Dive: the survey of 1,533 finance staff on how often these scams reach the people who can actually send money · Regula: a five country survey putting an average dollar figure on deepfake fraud losses · Pindrop: 1.2 billion analyzed calls and the rate at which fraud attempts now arrive
The Briefing arrives weekly: what moved, what it changes, what not to do, and a working tool for each. Every issue gives the tools away, free and built for the week they cover. No product tours, no victory laps. If it stops earning the read, unsubscribe.
This is back issue No. 008. Every tool in it still runs. Read the current issue · Open the archive
If something in here lands on a problem you are already working, and you would rather not work it alone, write to hello@goudegroup.com and tell us what you are exploring. You get a reply from a person, not a sequence.