Forty interviews in three weeks is achievable, but recruitment, not moderation, is the binding constraint in most programs. AI moderation removes the scheduling ceiling and matches human probe depth only on factual questions.
What actually limits interview throughput?
Almost always recruitment. A team commissioning forty interviews imagines the constraint is calendar capacity — how many conversations a moderator can hold in a day — and staffs accordingly. The actual constraint is how many people in the world match the screen and are willing to speak in the next fortnight.
That distinction changes what you do about it. Adding moderators to a recruitment-bound program adds cost and no speed. Loosening the screen adds speed and costs you the thing you were buying.
The diagnostic takes an afternoon. Count how many people plausibly match the screen in the target market, halve it for reachability, halve it again for willingness to speak inside the window, and compare what is left with the number in the brief. Where the arithmetic does not clear, the conversation to have is about the screen rather than about staffing, and it is much cheaper to have before the program starts than at the two-week mark.
How many interviews fit three weeks?
More than most teams assume, provided the screen is settled before the clock starts. The days lost at the front of a program are almost always spent deciding who counts as an expert, and that decision cannot be made in parallel with recruiting against it.
Where a screen is genuinely narrow — a specific role, at a specific scale, in a specific market, within a recency window — three weeks is tight and the honest answer is to scope a band rather than a number.
Where does AI moderation help most?
It removes the scheduling ceiling. An AI moderator can hold twelve conversations simultaneously at nine on a Tuesday evening, which is when a lot of senior operators are actually free, and it does so without a researcher's calendar becoming the bottleneck.
That helps precisely when recruitment is not the constraint: broad screens, large samples, known questions. On a narrow B2B screen it helps far less, because the people are the scarce thing and they were never waiting on a moderator's diary — which is the substance of choosing between an AI and a human moderator.
We hired two more moderators and the program finished on the same day it would have anyway. The people we needed were not sitting waiting for a call slot.
What does throughput cost in depth?
Consistency rises and depth falls. Running forty structured conversations produces forty comparable answers to the questions you asked. It produces fewer of the findings that come from a moderator noticing something nobody planned to ask about and following it.
That is a real trade rather than a failure. For a channel check or a pricing sweep, comparability is worth more than depth, and at that point an opinion survey across the same population often answers it more cheaply. For a discovery question where you do not yet know what matters, it is not.
How do you protect quality at scale?
Fix the discussion guide before the first call, code as you go rather than at the end, and reserve capacity for moderated follow-up calls. The programs that change a client's view are consistently the ones that kept room to go back to three people, and they start from a guide built around eight real questions rather than twenty.
Coding as you go matters most. A program that reaches interview thirty before anyone reads interview four has spent most of its budget asking questions that the early conversations had already answered.