Not Memory, Just Retrieval: Narrow Escape For OpenAI

  • Not Memory, Just Retrieval: Narrow Escape For OpenAI
    Listen to this Article

    In the first Indian decision to consider whether training a large language model on copyrighted material infringes copyright, Justice Amit Bansal of the Delhi High Court, in ANI Media Pvt. Ltd. v. OpenAI OpCo LLC, denied ANI Media's plea for an interim injunction against OpenAI on July 24, 2026. ANI, one of India's largest news agencies, raised both an output claim (that ChatGPT reproduced or misattributed its reporting) and an input claim (that OpenAI copied and stored its articles to train ChatGPT without authorisation), arguing that this fell within Section 14(a)(i) of the Copyright Act, 1957. OpenAI relied on the fair dealing exception for research under Section 52(1)(a), arguing that training an LLM is a transformative process rather than reproduction. OpenAI won on both counts. The ruling arrives as courts worldwide, from the United States to the United Kingdom, grapple with near-identical claims against AI developers, making it India's opening contribution to a global body of case law that is still taking shape.

    Jurisdiction

    OpenAI argued that the Copyright Act has no extra-territorial reach, since its models were trained in the United States. The Court held that although the Act is not extraterritorial, training on ANI's corpus and the resulting outputs formed one continuous process rather than two separate acts, and since ANI's principal place of business was in India and the outputs were accessed here, the Act applied. This lets an Indian court reach conduct that occurred substantially abroad, so long as some link in the chain touches Indian soil.

    When Retrieval Isn't Memory

    ANI's exhibits showed ChatGPT producing answers nearly identical to its news articles, suggesting the models had been trained on ANI's content. But the models named in the suit, GPT-4 and GPT-4o, were trained before the articles in question were published. What likely happened instead was that ChatGPT pulled current information from ANI's website while answering, retrieval-augmented generation (RAG), rather than recalling anything from training. The Court was direct: “the illustrations given in the plaint are post the training of Open AI's LLMs and a case for memorization… cannot be made out” (para 124, page 60).

    One question the Court left open: whether displaying a rightsholder's content through live retrieval amounts to “communication to the public,” a distinct right under Section 2(ff) of the Copyright Act. This right covers any work made available for the public to see, hear or enjoy, and does not require proof that the work was copied or stored, only that it was made accessible. The Court noted the issue but did not decide it, since ANI had not properly pleaded it; it will likely resurface. The underlying question is whether an LLM that retrieves and reproduces copyrighted content in real time, without ever storing it during training, can still amount to communication to the public under the present statute. That question is distinct from, and arguably harder than, the memorisation question the Court did resolve, since it does not depend on what happened during training at all.

    Where the Reasoning Strains

    The ruling is strongest on the output claim: the training-date mismatch was a clean, near-dispositive fact, and the Court's explanation of RAG is careful and well-informed. It is weaker in two respects. First, it conflates training and output when addressing jurisdiction, assuming the alleged infringement occurred wherever the model was trained. Second, it leans heavily on reading “research” in Section 52(1)(a) to cover AI training, even though India has no statutory exception for text and data mining, the use of automated tools to analyse large corpora, including for training AI models. Absent such an exception, the Court stretches “research” to accommodate AI training within the existing framework, doing by interpretation what Parliament has not yet done by amendment. This reasoning may only survive if higher courts, or the trial court itself, are willing to read the provision broadly.

    Three Doors the Trial Court Left Open

    The interim order decides only who wins the injunction application, not the case itself. Three issues remain open.

    First, what counts as “research” when the researcher is a machine. Section 52(1)(a) was written with a human reader in mind, but the Court applied it to AI training, which does not “read” as a human does, it processes entire corpora at a scale no person could. Whether the provision was meant only for human researchers, and not industrial-scale machine processing, remains unresolved. If courts eventually say it was, the training-side defence loses its footing, and OpenAI's win here would look far narrower in hindsight than it does today.

    Second, communication to the public. An LLM using RAG does not reproduce from stored knowledge; it serves live content in response to a prompt. Is that akin to a person reading and summarising a news article, or does it make the AI an unlicensed intermediary delivering the same information? ANI did not press this argument, and the Court did not examine it, but this is likely where future disputes will focus, since it bypasses the memorisation question and asks instead whether copyrighted material is being reproduced without a licence.

    Third, a procedural point. Interim relief in India turns on prima facie case, irreparable harm, and balance of convenience, a lower bar than a full trial on the merits. The Court's observations on the research exception and on how models like GPT-4 are trained were made without that fuller evidentiary record. ANI can still bring evidence at trial. The order may be cited by other litigants, but it is not a final position.

    What a Licensing Market Would Have to Look Like

    ANI did not seek a licensing remedy, but its underlying argument was that OpenAI should have paid to use its content. That connects to a wider debate over whether AI companies should pay for training data through collective licensing. Bartz v. Anthropic raised a similar point: Anthropic argued no real market existed for licensing data at the scale needed to train an LLM, given that companies would need millions of individual licenses and many online works have no identifiable owner. The Delhi High Court reached a similar conclusion: ANI could not show that a real licensing market existed for its content. In both cases, the absence of a working market cut against a finding of infringement, not in favour of one. That is a striking inversion: the very scale that makes licensing impractical is being read as a reason to excuse the copying, rather than as a reason to build the licensing infrastructure the market currently lacks.

    Collective licensing has a mixed track record, which counsels caution about relying on it here. In the United States, the Audio Home Recording Act's levy on blank tapes paid rightsholders only trivial amounts by 2020. The Authors Guild v. Google settlement, which would have created a body to license book digitisation, collapsed after objections from both rightsholders and users. Existing arrangements, such as those run by the Copyright Clearance Center or OpenAI's deal with Axel Springer, cover only a fraction of the material used to train major models, and none solve the problem of millions of rightsholders who are difficult to identify or trace.

    The European Union takes a stricter approach: rightsholders can opt out of having their works used for text and data mining, a mechanism India lacks. The EU AI Act goes further, requiring providers of general-purpose AI models sold in the EU to comply with EU copyright rules regardless of where training occurred. The Delhi judgment says nothing about OpenAI's liability under EU law, it holds only that ANI's works fall within India's exception for private use and research. The broader question of whether AI training must always be paid for remains open.

    Interim, But Not Inconsequential

    Justice Bansal was careful to note that the order does not settle the suit; its findings are preliminary. ANI can still present further evidence on memorisation and market harm, and the unresolved question of communication to the public could shape the outcome at trial. Even so, the order's two central holdings, that training can fall within the private-use research exception, and that outputs from live retrieval are not the same as memorisation, now stand as the first substantive precedent other Indian litigants will have to engage with.

    A working paper released in December 2025 by a committee constituted by the DPIIT proposed that AI companies pay for copyrighted training material through a mandatory blanket licence with a built-in default payment. That paper is not law, just as this interim order is not final.

    Will countries converge on whether copyrighted works can be used to train AI without permission or royalty? Probably not soon. As Professor Pamela Samuelson has written, “It would be beneficial to the growth of the generative AI technology industry and economies that depend on it if an international consensus emerged about the legal implications of using in-copyright works as training data for AI models” (p. 9). For now, that remains a wish, and until it materialises, courts like the Delhi High Court will keep deciding these disputes one statute, and one interim order, at a time.

    Author Krishna Ravi Srinivas is a DPIIT-IPR Chair Professor & Director, CoE in AI& Law at NALSAR University of Law, Hyderabad & Associate Faculty Fellow, CeRAI, IIT Madras. Feba Sara Vinu is a Research Assistant with IPR Chair at NALSAR University of Law, Hyderabad. Views are personal.


    Next Story