Start with the number nobody in hiring likes

In published testing, recruiters who believed they could spot AI-written text (67% claimed the ability) measured at roughly 52% accuracy — a coin flip wearing a confident face. Meanwhile, surveys consistently find that most applications now involve AI somewhere: drafting, tightening, or translating. Put those together and the common screening posture — “I can tell, and I reject what I catch” — fails on both halves: the telling is near chance, and the catching would disqualify most of the pipeline, including strong candidates.

For recruiters: what a detector can actually give you

A full-length cover letter usually clears the ~150-word floor below which no statistical detector is reliable, so a scan of one is meaningful in a way a scan of a two-line chat message never is. What an honest scan returns is a calibrated probability with a published error rate — not “AI DETECTED.” Read it the way you would read any screening signal with a known error rate:

  • A flag is a prompt to read closer, not to reject. At any realistic false-positive rate, screening hundreds of applications on a score alone guarantees you discard genuine writers.
  • The skew is documented. Research — including work from Cornell on who gets suspected — shows AI suspicion lands disproportionately on non-native English writers and formal stylists. An uncalibrated score operationalizes that bias; a calibrated one at least tells you how often it is wrong.
  • The letter's job has changed. If most applicants use AI, “did AI touch this?” stops separating candidates. What still separates them: specificity about your company and role, claims the candidate can defend live, and consistency with the interview. Detection can tell you a letter is statistically generic; only the interview tells you whether the person behind it is.

Practical policy: state your AI rules in the posting, scan only what is long enough to scan honestly, never auto-reject on a score, and weight detection below the two signals that actually predict performance — work samples and structured interviews.

For candidates: the honest view from the other chair

Using AI to draft a cover letter is not the integrity risk it is sometimes made out to be; submitting a letter that could have been sent to any company is the real failure, and it fails with human readers and detectors alike. Ground the letter in specifics no model could invent about you, follow any disclosure rule the employer states, and keep the final text something you can discuss fluently. If you want to know what an honest screener's tool would report, paste the letter into the free scan — you'll get the same calibrated probability a recruiter would see, with its error rate attached. And if a different tool's score ever costs you an opportunity you earned, the false-accusation playbook — including its workplace and freelance sections — is the response guide.

Freelance deliverables: agree before, verify after

The same logic covers Upwork- and Fiverr-style disputes from both directions. Clients: put your AI policy in the contract before work starts — “no AI” disputes after delivery, argued from a detector screenshot, are exactly the score-as-verdict failure this page warns about. Freelancers: a client's detector result is a probabilistic signal with a measurable error rate, not proof; your drafts, revision history, and process notes are stronger evidence than any percentage. Longer workplace-writing questions — including whether colleagues can tell day to day — are covered in the email guide.

What Cobalynx can and can't do here

A flagged letter doesn't mean a dishonest candidate, and a clean one doesn't mean a good hire. What we provide is a calibrated probability with published error rates — including the measured non-native false-positive rate — so nobody gets rejected on a coin flip dressed as a certainty. We can't tell you who typed a letter, we can't score texts too short to judge, and we won't certify anything as human — no tool honestly can. How the numbers are produced is on the methodology page; why detectors disagree with each other is in Can AI detectors be trusted?

Sources

  1. Published recruiter-perception testing: claimed detection ability (67%) versus measured accuracy (~52%) on AI-written application text (industry study, accessed Aug 2026).
  2. Cornell University research on bias in who gets suspected of AI use in professional writing contexts (news.cornell.edu, accessed Aug 2026).
  3. Liang et al., “GPT detectors are biased against non-native English writers,” Patterns (2023), arxiv.org/abs/2304.02819.

Common questions

Can recruiters really tell when a cover letter is AI-written?

Not by intuition: in published testing, recruiters who believed they could spot AI text (67% claimed to) measured at roughly 52% accuracy — coin-flip territory. Full-length letters are long enough for a statistical detector to do meaningfully better, but every detector carries a false-positive rate, so a flag is a signal to read closer, never a verdict.

Should I reject a candidate whose letter was flagged?

No — a flagged letter does not mean a dishonest candidate, and a clean one does not mean a good hire. Research also shows suspicion lands disproportionately on non-native English writers, which makes acting on an uncalibrated score a fairness problem as well as an accuracy problem. Use the letter's specificity and the interview, with any score as one weak input; our measured error rates are on the evidence page.

Is it bad to use AI for a cover letter?

Using AI assistance is now the norm rather than the exception; the differentiator is whether the letter says anything specific to you and the role. A generic AI letter fails because it is generic, not because it is AI. Follow the employer's disclosure rules where they exist, and make sure every claim in the letter is one you can defend in an interview.

Can I check what a recruiter's detector would see?

You can check what an honest one would see: paste your letter into the Cobalynx scan for a calibrated probability with a published error rate. Other tools are not calibrated to ours (or to each other), so no self-check guarantees another tool's verdict — if a screener's tool flags you unfairly, the false-accusation playbook covers the response.