Consent That Was Never Asked: India's Public Data Exemption In Age Of Foundation Models

Bhumika Singh

4 Aug 2026 10:00 AM IST

  • Consent That Was Never Asked: Indias Public Data Exemption In Age Of Foundation Models
    Listen to this Article

    In July 2026, Meta briefly rolled out Muse Image, an AI tool that let anyone generate images of a person built from that person's own public Instagram photos, without asking first. Every public account was opted in by default, and the opt-out was buried deep enough that most people never found it before the feature was pulled three days later. It's a small, self-contained episode, but it's a useful stand-in for a larger assumption: that once your data is public, using it for something else, including feeding it to a model, doesn't require asking you anything at all.

    Section 3(c)(ii) of India's Digital Personal Data Protection Act, 2023 rests on exactly that assumption, written cleanly into law. Publicly available personal data doesn't just get a lighter consent standard; it exists outside the Act's framework altogether. Purpose limitation, correction and access rights, security obligations, none of it applies once data clears the public-availability bar.

    This is a defensible starting position on its own terms. Posting something publicly does mean giving up the narrower claim that the post itself was private, and that principle has held up in privacy law for decades. But implied consent to being seen isn't the same thing as implied consent to being harvested. Someone scrolling through a profile or a journalist quoting a public post is limited by how much attention one person can give to a page at a time. A model trained on millions of such posts doesn't read them the way a person would. It compresses scattered fragments into inferences nobody volunteered, conclusions about someone's politics or habits that no single post gave away on its own. That compression doesn't expire either; it sits inside the model's weights indefinitely, and somewhere in the process, the data turns into a commercial input the person never had any say in. Piece together someone's scattered posts and you can end up with a fairly precise read on their health concerns, their politics, or where they live and when, details the person shared in fragments across months or years and never intended to see stitched into a single profile. No individual post crossed any line on its own. The assembly does. Section 3(c)(ii) treats both of these as the same act. They aren't.

    The EU has just answered this exact question, and answered it the other way. On July 7, 2026, the European Data Protection Board adopted Guidelines 03/2026 on web scraping for generative AI, confirming that the GDPR applies in full whenever scraping involves personal data, with no carve-out for AI training and no exception for content that happens to be public. A companion set of guidelines on anonymization lays out a three-part test for when data can genuinely be said to have left the GDPR's reach. The message is direct: public visibility settles whether someone can look. It says nothing about whether their data can be integrated into a training pipeline at scale, and India's law currently assumes it can.

    There's leverage India isn't using here, too. It has one of the largest, most active social media populations anywhere, and that scale of live discourse is exactly what Frontier Labs needs most and has the least of. The leverage is sharper than it first looks: frontier labs are increasingly short on well-annotated text outside English and Mandarin, and India generates enormous volumes of exactly that, Hindi, Tamil, Bengali, Marathi, and dozens of other languages, frequently mixed with English inside the same sentence, on some of the most active platforms anywhere. That kind of code-switched, multilingual corpus is hard to synthesize artificially and expensive to license on a per-piece basis. Being one of the few real sources of something usually means you can set terms for access to it. Right now, nobody is setting any.

    None of this requires repealing Section 3(c)(ii). A narrower fix would require three things, and there's already precedent for each elsewhere. Large-scale developers should have to publish a structured summary of training data sources. The EU's AI Act already enforces similar transparency requirements. Second, data principals should be able to register objections to specific public content used in training, administered through the Data Protection Board rather than left to platforms to police themselves. And third, exempt open, non-commercial, and academic development from the heavier obligations, so the rule constrains scale rather than entry.

    There's a related gap worth closing too. The DPDP Act treats anonymized data as outside its scope, but it nowhere defines what anonymization actually requires or how irreversible it needs to be. The EU's July 2026 guidelines fill that exact gap with a three-part test: whether a record can be isolated to a specific individual even without a name attached, whether records about the same person can be linked across datasets, and whether identity can be inferred indirectly from the data itself. A dataset only counts as anonymized if it fails all three. India has no equivalent standard, which means a developer can claim scraped or aggregated data was anonymized before training with nothing to hold that claim against.

    Transparency, not consent, is the right starting point. Most public-data training is probably an ordinary, non-displacing use, the digital equivalent of reading rather than republishing. The problem is that nobody can currently distinguish that from the narrower set of uses that would trouble even a permissive observer: content taken with no attempt to ask, or reused specifically to impersonate someone. Disclosure makes that line drawable; consent can follow once the conduct is visible. It cannot come before it, no matter how convenient that would be for everyone already building on top of public data.

    Author is a Law student at NLU Lucknow. Views are personal.

    Next Story