A CV used to be weak evidence that someone could do the job. It is now strong evidence that someone can produce a document. Those are not the same claim, and the gap between them is where a lot of hiring is currently going wrong.
The scale of the shift is not subtle. Applications per opening have roughly tripled since 2021 and now run past 300 for a single role. Independent analysis puts AI-generated text in around 78% of what arrives. On the employer side, 51% of organisations use AI somewhere in recruiting, up from 26% in 2024.
Both sides of the funnel got faster at producing text. Neither got better at producing evidence. The teams still hiring well in 2026 are not the ones reading faster. They are the ones who stopped treating the written application as the filter.
What a written application actually proves now
Screening was always a proxy exercise. Nobody believed a CV was the truth. The bet was that writing a decent one correlated with being decent at the work: you had to remember what you did, describe it clearly, and stay honest enough that a reference would not contradict you.
That correlation is what broke. Producing a polished, role-specific, keyword-aligned application now takes a few minutes. The skill it demonstrates is prompt-writing, and for most roles prompt-writing is not the job.
The practical effect shows up in the middle of the pile, not at the top. Strong applications look roughly as they always did. Weak ones no longer look weak. Everything that used to sit in the obvious-reject band, typos and generic cover letters and no mapping to the role, has been pulled up into a competent-looking middle. Your reject signal disappeared before your accept signal did.
That asymmetry is the part most teams miss. They read the flood as a volume problem and answer it with speed. It is a distribution problem. The spread between a strong candidate and a weak one narrowed on paper while staying exactly the same in reality.
Adding an automated screener makes the noise worse
The obvious response is to answer machine text with machine reading. Around 90% of incoming applications now pass through some automated ranking before a person sees them, and that is sold as the fix for volume. It has two failure modes worth knowing before you buy one.
The first is selection bias pointing the wrong way. Automated screeners score text, and text produced by a language model is unusually well-formed for exactly the features those screeners reward: clean structure, dense keyword coverage, explicit outcome statements. Testing has put self-preference rates above 80%, meaning the screener favours AI-written applications over human-written ones. You installed the filter to catch machine text and it promotes it.
The second is escalation. Filters tighten, candidates add tooling to clear them, filters tighten again. Every round costs both sides more and removes information from the middle. Nobody wins a race where both sides can automate their next move.
There is also a regulatory tail that European buyers have not all priced in. Recruitment and candidate-ranking systems are classed as high-risk under Annex III of the EU AI Act, which attaches documented risk management, bias testing, record-keeping, a named human with genuine authority to override the output, and an obligation to inform worker representatives before deployment. It applies to any organisation whose AI outputs affect people in the EU, wherever that organisation sits. The timing moved: the Digital Omnibus, approved by the Council on 29 June 2026, deferred those Annex III obligations from 2 August 2026 to 2 December 2027. The transparency rules and the AI literacy duty did not move. A deferral buys preparation time, not an exemption.
Volume is not your problem. Evidence is.
The two arrive together, which is why they get confused, but they have different fixes and only one of them is expensive.
Volume is an operational load. Three hundred applications is a throughput question, and throughput is solved with process, capacity, or a partner. Unpleasant, but bounded.
Evidence is a different question, and throughput does nothing for it. If every application makes the same claims to the same standard, reading all 300 gives you what reading 30 gave you: a ranking of writing quality.
You can tell which problem you have by looking at what happens after the shortlist. If shortlists arrive on time and interviews keep surfacing people who cannot do the work, your evidence is broken, and more screening capacity will not touch it. If shortlists are late but the people on them are right, that is volume, and volume is fixable with hands. Buying a volume solution for an evidence problem is the most common and most expensive mistake in this whole area.
Three things that still count as evidence
What survives is anything a candidate has to do rather than describe.
A work sample drawn from the actual job. Not a puzzle, and not a take-home that costs someone a weekend. A short realistic task: rewrite this angry customer reply, triage these five tickets in order and say why, fix this function. Thirty to sixty minutes, graded blind against a rubric written before any answers arrive. A model can help a candidate produce it, and that is fine. If the task is close enough to the work, using a model well is part of the work. What you are grading is judgment, and judgment does not survive being handed to a prompt.
A structured reference about one specific incident. General reference calls are theatre. A useful one names a single event: tell me about the hardest week this person had on your team, what happened, what did they do, what would you have wanted done differently. One concrete story from a former manager carries more information than a page of self-description, and nobody can generate it on your behalf.
Evidence that already exists in the world. A repository with real commit history, a support account with a review trail, a portfolio, a community contribution. It carries a timestamp and a witness. That is what makes it evidence rather than a claim.
| What you look at | What it actually proves | What it costs you | Effect of AI tools |
|---|---|---|---|
| The written application | The candidate can produce a document | Minutes each, days in aggregate | Proves almost nothing now |
| An automated keyword or fit score | The text matched a pattern | Low, plus licence fees | Prefers machine-written text, above 80% |
| A short work sample from the real job | Judgment on a task you recognise | Design once, 10-15 min to grade each | Holds, if the task mirrors the work |
| A structured reference on one incident | Behaviour under real pressure | About 20 minutes per call | Cannot be generated |
| Public work with a history | Sustained output over time | Minutes to verify | Cannot be generated |
Look at what those three have in common. Every one of them is something a candidate could not manufacture in the ten minutes before applying.
The screening work does not disappear. It moves.
This is the part that decides whether any of the above actually happens.
Reading 300 applications is bad work, but it is easy to start. It needs no preparation, no rubric, no decisions in advance. Running a work sample needs someone to write the task, agree what a good answer looks like before answers exist, grade consistently, and hold the line when a likeable candidate scores badly. That is less total work and much harder to begin.
So the effort migrates from reading to designing, and design is where hiring processes stall. It also has a prerequisite: you cannot write a gradeable task until you know what the person is actually for. A vague requirement list produces a vague task and an ungradeable answer, which is the same failure that produces rejected shortlists in the first place.
The sequence is worth stating plainly. Define the role precisely enough that you could grade an answer to it. Build the task from that definition. Then open the funnel. Teams that open the funnel first end up doing the only thing available to them, which is reading. And the cost of getting it wrong is documented: a bad hire on a small team is a good deal more expensive than the 30% of salary usually quoted.
Who does this work when you cannot
Three hundred applications plus a rubric-graded work sample is real work, and in a company of 10 to 100 people it usually lands on somebody who already has a full job. There are three honest ways to handle that.
Run it in-house. This works when hiring is continuous enough for the practice to mature: a recruiter or an operations lead who runs the same process often enough to grade consistently. Below roughly one hire a month the practice never matures, and every search restarts from nothing.
Bring in a recruiting partner for the search. The partner absorbs sourcing, the first pass, and the calibration loop, then hands over a shortlist you interview. That is our recruiting line: we work the Ukrainian and wider Eastern-European market, and internationally, and we charge a flat fee, paid when someone is hired. For operations and support roles that is a low four-figure fee, and it does not scale with the salary, so nobody on our side has a reason to talk the number up.
Have the function run for you as a managed team. This is a separate option, not a variation on the one above. Instead of hiring people onto your payroll, an outsourced team runs the function as a fixed monthly team, not per ticket. It removes the hiring question rather than answering it, and it fits when the work is continuous and you do not need the headcount on your own books.
Keep the two apart when you compare them. A placement puts a person on your payroll and is paid once. A managed team does not and is not. They are priced differently because they solve different problems.
How IMMIDO screens
We have run recruiting since 2022, alongside 24/7 first-line support, worked inside your own tools. We also built and ran our own support pod for a regulated operator for more than five years, seven agents, so none of this is theoretical. We have run the seat, and we screen for what it actually takes to be good in it.
Which is why we do not shortlist on written applications. They set the longlist and nothing more. We screen against the real role rather than the resume, testing for the work a person will actually do in week one instead of the words on the page, and that grading happens before anyone reaches your shortlist.
A first shortlist typically lands in 7-10 business days. Not because we read faster, but because we source from an existing pipeline and the grading is defined before the applications arrive. That is also why the recruiter-controlled half of time to hire is the half that responds to process at all.
If you are staring at a folder of applications that all look the same, the fix is not another pass through them.