Click the Play button to listen to article

When investigators ask a suspect to read a script into a recorder so a forensic laboratory can match the voice against an intercepted call or a ransom recording, the exercise looks, on its face, like routine identification no different from taking a fingerprint or a photograph. Yet unlike a fingerprint, a voice sample is produced by an act indistinguishable, in form, from speech itself. That formal similarity is what makes voice-sample compulsion one of the more unsettled corners of Article 20(3) jurisprudence, and it deserves sharper scrutiny than the Supreme Court's treatment of the issue in Ritesh Sinha v. State of U.P. has so far given it.

The testimonial/physical divide, and its strain

Article 20(3) protects a person from being “compelled to be a witness against himself.” The Supreme Court's foundational reading of that clause, in State of Bombay v. Kathi Kalu Oghad (1961), drew a line between furnishing evidence that is “testimonial” communicating the contents of one's own knowledge and furnishing evidence that is merely physical: specimen handwriting, signatures, thumb impressions, blood, or hair. Only the former attracts the protection. This distinction survived, and was substantially reaffirmed, in Selvi v. State of Karnataka (2010), where the Court struck down involuntary narco-analysis, polygraph tests, and Brain Electrical Activation Profiling precisely because they extracted testimonial content from the subject's mind under compulsion, regardless of whether the resulting statements were later used as formal confessions.

Voice sits awkwardly across this line. Its acoustic properties pitch, formant frequencies, vocal tract resonance are, in one sense, as physical and involuntary as a fingerprint ridge pattern; a person cannot alter them at will, and comparing them to a disputed recording is conceptually similar to comparing two ridge patterns. But producing a voice sample also requires the subject to speak, and speech is the very medium in which testimonial self-incrimination occurs. A fingerprint carries no propositional content; a voice sample, especially one built on a script chosen by the investigating officer, can be made to carry exactly the content the prosecution needs down to the specific words allegedly used during the offence.

Ritesh Sinha: a closer look

The case that settled the threshold question had an unusually divided journey. It arose from a 2009 FIR alleging that the appellant, in conspiracy with a bank manager, had extorted money from the complainant; the prosecution sought to compare a recorded conversation with the appellant's voice, and a magistrate in Uttar Pradesh directed him to furnish a sample. The appellant resisted on two grounds that no CrPC provision empowered a magistrate to compel a voice sample, and that doing so would violate Article 20(3).

When the matter first reached the Supreme Court in 2012, a two-judge bench split. Ranjana Prakash Desai J. held that, absent express statutory authorisation, a magistrate had no such power Parliament had amended the CrPC in 2005 to insert Section 311-A, empowering magistrates to direct specimen signatures and handwriting, but had conspicuously not extended it to voice, and that omission could not be filled by interpretation. Ranjan Gogoi J., dissenting, held that magistrates could draw on inherent powers pending legislative correction, treating the gap as an oversight rather than a deliberate exclusion. The split led to a reference that sat before a larger bench for seven years.

The reference was resolved in 2019 by a three-judge bench Chief Justice Gogoi, writing for himself and Deepak Gupta and Sanjiv Khanna JJ. which adopted his own 2012 dissent as the law. The Court read into the CrPC, by analogy with Sections 53, 53A (medical examination of an accused) and 311-A, an implied magisterial power to order voice samples, while asking Parliament to legislate and remove the ambiguity. On Article 20(3), it relied squarely on Kathi Kalu Oghad and Selvi, holding that a voice sample, like a signature or fingerprint exemplar, establishes identity through comparison alone and communicates no personal knowledge of facts in issue.

Two features of this reasoning mark where later scrutiny needs to begin. First, the Court treated the Article 20(3) objection almost as a corollary of the structural question whether the power to compel a sample exists at all borrowing the Kathi Kalu Oghad categorisation wholesale rather than testing voice against it independently. Second, its invitation to Parliament was never taken up in a way that addressed content selection. The Criminal Procedure (Identification) Act, 2022, which followed three years later, dealt with retention periods, categories of measurements, and offence-seriousness thresholds, but left unanswered the question Ritesh Sinha never asked: what the accused may be made to say while the sample is taken.

Where identification quietly becomes reconstruction

This is not hypothetical. Investigating agencies routinely ask suspects to repeat, for matching purposes, the specific words captured in an intercepted call or threatening message “give me the money by tomorrow,” say, if that phrase forms part of the alleged extortion recording. The stated purpose is legitimate: matching is easier when the comparison sample uses the same phonemes as the disputed recording. But the effect is that the accused is made, under compulsion, to utter words mirroring the alleged crime, in a manner a fact-finder could plausibly treat as corroborative of authorship an outcome close to compelling an admission, dressed in the vocabulary of forensic science.

This is where the Kathi Kalu Oghad framework, built for handwriting and fingerprints, begins to fail voice as a category. A signature exemplar does not require the writer to inscribe the specific defamatory sentence under investigation; a fingerprint does not require the subject to press it in the shape of the crime-scene mark. Voice-sample protocols, by contrast, often do exactly this selecting content for its evidentiary resemblance to the offence rather than merely its phonetic utility. Identification and reconstruction collapse into a single compelled act.

A content-neutrality test as the missing safeguard

The doctrinal fix is not to abandon the physical/testimonial distinction it remains sound for fingerprints, DNA, and iris scans, which inherently involve no propositional content. The fix is to recognise that voice occupies a hybrid category and import a safeguard the Court has not yet articulated: a content-neutrality requirement. A compelled voice sample should be treated as non-testimonial only where the script is neutral standardised, unrelated in wording to the facts in issue, chosen for phonetic coverage rather than evidentiary resemblance. Where investigators instead direct the subject to reproduce case-specific language, the exercise shifts in character, becoming functionally closer to a coerced reconstruction of the offence, and ought to attract the scrutiny Selvi applied to narco-analysis.

This has a practical anchor in identification-parade jurisprudence, where courts have long been alerted to suggestion contaminating an identification exercise; a parade conducted with cues pointing to the suspect is treated as unreliable and coercive in substance even if procedurally compliant. Voice-sample protocols deserve the same lens a facially neutral “physical evidence” exercise can be substantively transformed into a testimonial one by the content investigators choose to elicit.

The statutory gap

The 2022 Act's definition of “measurements” extends to biological samples and “behavioural attributes including signatures, handwriting,” and could plausibly be read to cover voice yet the statute never names voice samples explicitly, and Parliament's Statement of Objects and Reasons is silent on the point. This is the very gap Ritesh Sinha asked Parliament to close; courts are still filling it through inference rather than clear statutory text, including on questions the Act does address for other measurements, such as retention limits and authorisation.

That gap matters because it is precisely the space where content-neutrality safeguards would need to live: script standardisation, magisterial pre-approval of the words to be used, and a clear rule that the sample cannot itself be led as evidence of what the accused said in connection with the offence, only as a comparator for voice-print matching. Absent that, Article 20(3) risks becoming a right that protects the fingerprint but not the sentence a distinction the framers, addressing "witness against himself" in unmistakably testimonial language, are unlikely to have intended.

Ritesh Sinha resolved the threshold question of whether voice samples may be compelled at all. It did not resolve because it was not squarely asked to the harder question of how such samples must be obtained so that identification does not shade into incrimination. Until Parliament or the Court supplies a content-neutrality safeguard, the line between forensic voice matching and compelled self-incrimination will continue to be drawn, case by case, by the investigating officer choosing the script.

Author Raghvendra Kumar Chaudhary is an Assistant Professor at CHRIST (Deemed to be University), Delhi NCR Campus & Abhijit Mishra is currently working as an Assistant Professor at the School of Law, Bennett University. Views are personal.

Tags: