How a Profitable Retainer Quietly Becomes an Unprofitable One

The two costs that kill a retainer margin without appearing on any scope document — and the re-scope conversation that saves the account.

Author
Prabhash Jha
Published
Reading time
14 min read

The retainer was profitable in month one. It was, in the polite phrase agencies use to describe the moment the finance meeting stopped being enjoyable, “under review” by month six. Nothing in the scope had changed on paper. The fee was the same. The deliverables were the same. And yet the account was quietly losing money every week, and the founder who ran it did not notice until the P&L for the quarter was already printed.

This post is about how that happens, and — more usefully — the two costs that cause it that never appear on any scope document, the monthly measurement that catches them before the P&L does, and the re-scope conversation that saves the account instead of ending it.

The advice that misses

If you type “how to prevent a retainer becoming unprofitable” into a search engine, the answer you get is roughly the same across page one. Track hours per client. Flag the account when time-spent reaches 75% of the budgeted hours. Have a monthly reconciliation call. Charge for out-of-scope work.

All of this advice is fine. It is aimed at a specific failure mode — scope creep, in the visible sense that a client asks for an extra landing page and you build one — and it works for that failure mode. What it does not catch is the failure mode that most retainers actually die from, which is not new tasks the scope document did not cover. It is old tasks that the scope document does not describe accurately, done more expensively than the fee assumed, in ways that will never show up on a timesheet because nobody thinks to log them.

The two costs that do the killing are response latency and unbilled thinking time. Both of them are invisible to the industry-standard tools, both compound weekly, and both are covered by the same re-scope conversation script — the one below — which no vendor is going to publish because the vendor sells the dashboard that does not see them.

Cost one: response latency

Response latency is the cost of a client whose questions get answered too quickly.

Every retainer starts with an unspoken assumption about how fast the client can expect to hear back from you. If you never make it explicit, the client sets the default themselves, and the default they set is: whenever they email, whenever they message, whenever they call. In month one, when there are three other clients and the account is fresh, it is not expensive to reply inside the hour. By month six, when there are eleven other clients and this one has grown a habit of pinging you at 4pm on a Wednesday about a screenshot they took at 3:55pm, the same hour-inside response is expensive in a very specific way.

The way it is expensive is that it fragments the working day of your most-billable people. A senior producer who could have done four hours of concentrated work in an afternoon can instead do six twenty-minute stretches broken up by two-minute Slack replies, and the difference between those two afternoons is the difference between an account that runs on margin and an account that runs on goodwill. The client has not asked for anything the scope document does not describe; they have asked for a screenshot in a Slack message. The cost of the reply is not the two minutes it took to send it. It is the twenty minutes on either side of it, which is the well-documented cost of interrupted work.

Almost no agency measures this. Not because they do not know it is happening — the producers know exactly when it is happening — but because the tooling to measure it does not exist and, honestly, the client would find the measurement embarrassing. So the cost stays invisible, and the account owner reports on hours logged and deliverables shipped and the margin drips out of a hole nobody has named.

Cost two: unbilled thinking time

Unbilled thinking time is the cost of the work that gets done before the work you are billing for.

Every reasonable retainer includes deliverables — a monthly report, a set of creatives, a campaign build, a strategy note. The scope document is written around those deliverables. What the scope document does not say is that most of them take about half as long to produce as they take to think through, and the thinking-through is what actually consumes the account manager’s brain.

Consider the “monthly performance review” that appears on almost every marketing retainer. The document itself, once you know what you want to say, takes maybe two hours to write. The thinking-through — reading the last four weeks of dashboards, running the diagnostics, working out what actually changed and what is noise, framing it in a way the client will trust — takes six. Nobody logs those six hours anywhere. On the timesheet, the account is booked for “monthly report — 2 hours”. On the payroll, the account manager has spent Tuesday, Wednesday morning, and part of Thursday morning thinking about this client. The retainer was priced for the timesheet. The account is being delivered by the payroll.

Multiply that gap across ten clients and it is not a gap. It is the entire margin.

The compound

The two costs are not additive. They are multiplicative.

A client who has trained you to respond within the hour is also a client whose account manager cannot get any thinking time done during the working day. So the thinking-time cost — the six unlogged hours per month — moves to evenings and weekends, where it is done tired, badly, and by people who will eventually resign. The retainer looks fine on paper for another quarter. The turnover in the team is where the cost actually lands.

Meanwhile, the client’s own perception of value has been quietly recalibrated by the response latency. A retainer whose deliverables are shipped monthly but whose replies arrive within the hour trains the client to see the retainer as an on-call resource. They stop remembering the deliverables and start remembering the availability. When it is time to renew, the negotiation is not about the deliverables you have shipped; it is about the availability they have come to expect. The next month’s fee is priced against the deliverables, and the next month’s cost is delivered against the availability.

The monthly measurement that actually catches this

You do not need a timesheet. The people who fill in timesheets have been trained to fill them in against the scope document, so timesheets tell you exactly what the scope document already told you.

What you need is a two-column note per client, done in twenty minutes at month-end by whoever owns the account, and it looks like this.

Column one: shipped this month. The list of things that came off the account this month. Deliverables, reports, campaigns, strategy notes. Anything the client received. Written in plain language, one line each.

Column two: everything else this month. Everything the account team did that did not appear in column one. Every meeting. Every Slack conversation over ten messages long. Every email chain over five replies. Every ad-hoc analysis. Every time somebody redid a piece of work because the first version got sent back. Every 4pm-on-Wednesday screenshot request. Written in plain language, one line each.

You are not measuring hours. You are counting entries. If column two has more entries than column one, the retainer is running on unbilled thinking time and response latency, and the numbers on the P&L are lagging what has already happened.

The reason this measurement works when timesheets do not is that it does not ask anyone to remember how long something took. It only asks whether it happened. That is a question a human being can honestly answer at month-end; how long something took is a question they cannot.

The three warning signs before the P&L catches up

If you do the two-column note for three months, the pattern will announce itself before the finance meeting does. The warnings, in the order they appear:

Sign one: the account manager stops proposing anything. In month one, the account team was full of “we should try X”. By month six, when the same team is asked what to try next, they say “we should keep doing what we’re doing”. This is not agreement with the strategy; it is the operational reality that they do not have any thinking-time left. A retainer whose team has stopped proposing new work is a retainer that is delivering only from muscle memory, which is the most expensive way to deliver anything.

Sign two: the client’s Slack messages start with “quick one”. The word “quick” in a client message is almost never accurate. What “quick one” actually means is: this is a small enough thing that I do not feel bad asking, but a specific enough thing that I need a real answer. The client is not being manipulative. They have simply learned that “quick one” gets a response inside the hour. When the pattern establishes itself, you have — without noticing — become a shared inbox for their operations.

Sign three: the deliverables start slipping and nobody knows why. The monthly report is a day late, then two, then three. The team is not idle; the team is exhausted. Nobody is doing less work; they are doing less work that is on the scope document, because everything not on the scope document has to happen first, in real time, at 4pm on a Wednesday.

When you see all three signs on the same account, the P&L reconciliation for that account is going to be ugly. You have between four and eight weeks. Have the conversation now.

The re-scope conversation

The reason this account is losing money is not that the fee is wrong; it is that the delivery model is wrong. The conversation you need to have is the one that fixes the delivery model, and the reason most agencies never have it is that they think the conversation is a request for more money. It is not. It is a request for the account to survive.

Here is the shape of the conversation, in three moves. Each move is one sentence long, in normal working language, said by the account owner and heard by whoever holds the budget on the client side. This is not a script to be recited; it is the sequence and the specificity.

Move one: name the pattern, not the client’s behaviour. The account owner opens with a sentence like: “This retainer was built around monthly deliverables, but the work has moved toward real-time back-and-forth, and I want to reset how we run it before the next quarter, so the delivery model matches how we’re actually using each other.” Note what this does not say. It does not blame the client for asking too much. It does not say the retainer is unprofitable. It names the pattern and asks to re-scope, both of which are neutral facts.

Move two: propose the two structural changes, in a form the client can accept without losing anything. The two changes are almost always the same. First: response windows. Slack replies within one working day rather than one hour, urgent things flagged as urgent, and one weekly office-hours block for real-time questions. Second: strategy time. A named half-day per month for the account team to think through what is next, protected from the ongoing work. Neither change costs the client anything in absolute terms; both save the account.

Move three: offer the trade. Where a fee increase is warranted, this is the point at which you name it, and you name it in exchange for something specific — a new deliverable, an expanded reporting cadence, a stated ambition. Where a fee increase is not warranted, this is the point at which you propose the delivery-model reset without a fee change and offer to review the fee in the following quarter if the model has settled. The important thing is that the trade is specific. “We need to raise the fee” is a fight. “We are moving to a model where we can protect strategy time; here is what that adds” is a conversation.

Have the conversation with whoever holds the budget, in person or on video, never by email. Fifteen minutes on a calendar invite booked two days in advance is enough. If you cannot get the meeting, you have already found out something about the account that is worth acting on separately.

What to do if the conversation fails

Some accounts will not accept the re-scope. When that happens, you have three options, and the one you pick is a business decision, not an operational one.

Option one: raise the fee to match the delivery model. If the client wants to keep responding in the hour and does not want a named half-day for strategy, the price for that is a fee that assumes the higher cost of delivery. Model it honestly — if the current fee is X and the account is being delivered at cost 1.3X, the fee that makes it profitable at the current delivery model is closer to 1.6X than 1.3X, because you also need margin. Present the number with the reasoning; be prepared for the client to leave. That is a fair outcome.

Option two: reduce the deliverables to match the fee. If the fee cannot go up and the delivery model cannot change, the honest response is to reduce what you ship each month so the account’s actual cost matches the actual revenue. Drop the monthly report to a quarterly one. Drop the strategy note. Whatever you drop, name it, and get the client’s written acknowledgement that this is now the scope. This is unpopular and it is honest.

Option three: end the retainer at renewal. If neither of the above works, the account is unprofitable at any delivery model the client is willing to accept, and the right move is to let it end at renewal, on good terms, with a handover that protects the client. Do not fire the client mid-quarter; time it to a natural boundary and offer to introduce them to two competitors who could serve them well. This costs you nothing that was not already lost, and it saves the reputation on both sides.

The one option that is not on the list is “continue and hope”. The account is not going to fix itself. Every month you do not act, the two-column note gets more lopsided and the account manager gets closer to a resignation letter.

Why this is not a dashboard problem

The temptation, having read the above, is to buy a piece of software that measures response latency and thinking time. There are many such pieces of software. They will not solve the problem, for two reasons.

The first is that measurement without a conversation is just data. If the account owner sees a red number on a dashboard once a month and does not have a script for what to do about it, the red number becomes background — the sort of thing you learn to notice without acting on. The two-column note works because it is done by the person who has authority to act, in the same twenty minutes it is done in. There is no lag between measurement and action.

The second reason is that the thing that fixes the account is not the data. It is the willingness to have the conversation. Every account owner who has ever run a retainer knows, without a dashboard, which of their accounts are running on goodwill. The measurement is not what surfaces it; the measurement is what gives them cover to raise it. If you are the founder, your job is not to buy them the measurement. It is to give them cover, and to have the conversation yourself on the accounts where they cannot.

Where this fits with the rest of the work

The retainer economics conversation sits next to a small number of related conversations, and understanding all of them together is more useful than understanding any of them alone.

If your problem is that you can see the account is unprofitable but you cannot see the cash consequence yet, the difference between what you are measuring and what your bank balance is measuring is covered in Cash flow vs profit: the difference that sinks most small businesses. A retainer can be loss-making for six months before the cash-flow statement notices, and the reverse is also true. Do not conflate the two.

If your problem is that you have not yet built a retainer at all and you are pricing services one project at a time, the framework for setting a rate that protects against exactly the failure mode above is in How to price your services as a freelancer or consultant. Pricing badly at the start makes the re-scope conversation harder for years.

If your problem is that a client has already stopped paying — not is-not-quite-profitable, but is-not-paying-at-all — the response is a different sequence and a different set of levers, in The client has stopped paying: here is the sequence, in order.

The one-sentence version

A retainer becomes unprofitable when the delivery model outgrows the scope document. The delivery model outgrows the scope document by two things — response latency and unbilled thinking time — and both are invisible to the tools that measure hours. The re-scope conversation is what turns the invisible into a decision, and the decision is what saves the account.

Do the two-column note this month. Have the conversation before you present the P&L. It is a shorter meeting than the one where you tell the team you are letting an account go.

Rules of thumb are a poor substitute for your own figures. Work out your emergency fund target.

Keep reading