IV The Human System
What Deception Detection Actually Detects
A validated marker is a small average difference between two populations under stated conditions, which is a different object from a tell — and the gap between them is where an entire market makes its living.
A discipline is worth what it is prepared to refuse. So the useful place to start is not with what this house rejects but with the standard it holds itself to, because scientifically defensible is a phrase that can be said at length without meaning anything.
It means four commitments. That a finding is published where people with a professional interest in demolishing it can reach it. That the effect has been measured rather than asserted, with a size and a rate at which it fails. That the technique was built and tested by practitioners whose institutional record is checkable — intelligence and national security work, behavioural psychology, neuroscience. And that a laboratory result has been carried into field conditions and survived. The fourth is where almost all the trouble lies, and it is the one the surrounding trade skips.
A marker is a population statistic
The most heavily cited work on behavioural cues to deception pooled 1,338 estimates of 158 cues from 120 independent samples, and appeared in Psychological Bulletin in 2003. It is often summarised as though it vindicated the search for tells. It did close to the opposite. Liars, the authors report, are in some respects less forthcoming, tell less compelling tales, make a more negative impression and are more tense — and, in the same abstract, many behaviours showed no discernible link, or only a weak one, to deceit.
Magnitude decides this, not direction. Of the 88 cues supported by at least three independent estimates, 23 produced a combined effect larger than 0.20 in absolute terms and 65 sat at or below it, 0.20 being the conventional marker of a difference a person can barely perceive. The cues also grew stronger when the speaker was motivated, and stronger again when identity rather than money was at stake. The signal depends on the conditions, not only on the lie.
That is what a validated marker is: a small average difference between two populations, established under stated conditions, with a documented habit of vanishing when the conditions move. It tells an investigator where to look. It is not a property of the person sitting opposite, and it does not survive the step from an average to an individual.
Which is why the trained version fails too
The largest field test of the contrary proposition was run by a government and the results were published. By 2013 the United States Transportation Security Administration had spent roughly $900 million screening passengers by observation of behavioural indicators. The Government Accountability Office reviewed four meta-analyses covering more than 400 studies from the preceding sixty years and reported, on 8 November 2013, that human ability to identify deceptive behaviour from behavioural indicators is the same as or slightly better than chance.
Revision followed rather than retirement, which produced a sharper figure. On 20 July 2017 the GAO examined the 178 sources cited in support of a shortened list of 36 indicators. Three of them qualified as valid evidence for the indicator they were cited for. Eight of the 36 indicators had any valid support at all.
That is not a finding about inattentive observers. It is a finding that the signal is too faint to carry a judgement about one person, however well the watcher has been taught.
The instruments that promise a verdict
Better instrumentation has not closed the gap, and the official assessments are blunt about why. The National Research Council's 2003 review of the polygraph found that the physiological responses the instrument measures are not uniquely related to deception, that its theoretical rationale is weak, and that the evidence overall is scanty and scientifically weak. Its most generous conclusion was that specific-incident tests can discriminate lying from truth-telling well above chance though well below perfection, among examinees untrained in countermeasures, with no precise estimate of accuracy available.
Imaging is the newer promise and has been tested in a harder room. Research in functional magnetic resonance imaging does allow the brain's response to stimuli and stressors to be observed as it occurs, and this house says so on its own page. Whether it delivers a verdict on one person's particular answer is a separate question, and in September 2012 the United States Court of Appeals for the Sixth Circuit reached it as a matter of first impression in any jurisdiction. Excluding fMRI lie-detection testimony was within the trial court's discretion under Rule 702, the court held, because the technology had not been fully examined in real-world settings and the testing performed had departed from the research protocols. Beneath that sat a starker finding: no known error rates outside the laboratory. Watching a brain respond is not adjudicating a claim.
Why the market never closes
The demand is older than any of these instruments and will not be argued away. Nobody wishes to be deceived, and not knowing is a continuous cost people pay to be rid of. Pseudoscience is not principally a supply of bad evidence; it is a supply of certainty, and on that one attribute the evidence base cannot compete.
Congress recognised the shape of that demand in 1988. The Employee Polygraph Protection Act, approved on 27 June 1988 and effective six months later, made it unlawful for most private employers to require, request, suggest or cause an employee or applicant to take a lie detector test, or to use or inquire about the results. The definition is the instructive part. It reaches a polygraph, a deceptograph, a voice stress analyser, a psychological stress evaluator, or any similar device used for rendering a diagnostic opinion regarding the honesty or dishonesty of an individual. Congress defined the category by the promise rather than the mechanism — and the promise outlived the statute, migrating to methods that make the same offer without a machine to regulate.
Neuro Linguistic Programming and the frameworks built on body-language interpretation are where much of it went. The honest statement of the evidence is narrower than the mockery it attracts, and the narrowness is the entire standard. A systematic review in the British Journal of General Practice in 2012 found ten experimental studies of NLP against health outcomes, judged the risk of bias high or uncertain across all of them, and concluded that there is little evidence of benefit — while stating expressly that this reflects the limited quantity and quality of the research rather than robust evidence of no effect. Unproven is a different verdict from disproved, and a house that will not overstate its own findings does not get to overstate that one. These frameworks have not met the standard, and they are sold as though they had.
Elicitation produces information; detection produces a verdict
The state that mandates the instrument has already legislated the distinction. Section 28 of the Offender Management Act 2007 permits the Secretary of State to impose a polygraph condition on the licence of certain offenders released at eighteen or over. Section 30, in force since 19 January 2009, then withholds the verdict: no statement made by the released person while participating in a polygraph session, and no physiological reaction recorded while they were questioned, may be used in proceedings against them for an offence. The session is kept for what is said around it and discarded as proof of anything.
The international standard points the same way. The Principles on Effective Interviewing for Investigations and Information Gathering, adopted in May 2021, set the object of an interview as accurate and reliable information rather than confirmation of what the interviewer already suspects, on the premise that coercion is efficient at producing statements and poor at producing true ones.
So, the title. Deception detection does not detect lies. It detects distance — between an account and the documentary record, between one telling and the next, between what a person volunteers and what has to be drawn out, between the effort a claim appears to cost the speaker and the effort it should cost if it were remembered rather than constructed. Each is checkable against something outside the conversation, which is the only reason any of it is worth commissioning. The product is a ranked list of things to establish and the observation that would settle each. Records close the question. The interview only opens it properly.
The studies, statutes and judgments cited here are public instruments, set out so a reader can check them. Nothing here is legal advice, and no assessment of any person or company is offered or implied.
The method sits under Deception Detection, Truth Elicitation & Human Behavior; the commercial application, where an account is tested against a record, under Due Diligence Investigations; the population-level version of the same science under Societal Research & Operational Human Terrain.
Sources
- DePaulo, Lindsay, Malone, Muhlenbruck, Charlton & Cooper, "Cues to Deception", Psychological Bulletin (2003) 129(1) 74–118
- GAO-14-159, Aviation Security: TSA Should Limit Future Funding for Behavior Detection Activities (US Government Accountability Office, 8 November 2013)
- GAO-17-608R, Aviation Security: TSA Does Not Have Valid Evidence Supporting Most of the Revised Behavioral Indicators Used in Its Behavior Detection Activities (20 July 2017)
- National Research Council, The Polygraph and Lie Detection (National Academies Press, 2003)
- United States v. Semrau, 693 F.3d 510 (6th Cir., 7 September 2012), No. 11-5396 — opinion of the court
- Employee Polygraph Protection Act of 1988, Public Law 100-347, 102 Stat. 646 (Statutes at Large, govinfo)
- Offender Management Act 2007, section 28 (application of polygraph condition)
- Offender Management Act 2007, section 30 (use in criminal proceedings of evidence from polygraph sessions)
- Sturt, Ali, Robertson, Metcalfe, Grove, Bourne & Bridle, "Neurolinguistic programming: a systematic review of the effects on health outcomes", British Journal of General Practice (2012) 62(604) e757–e764
- Principles on Effective Interviewing for Investigations and Information Gathering (the Méndez Principles), May 2021
- Privy Consul — Deception Detection, Truth Elicitation & Human Behavior