Key takeaways

  • In an interim ruling, the Delhi High Court rejected a request by Indian news agency Asian News International (ANI) for a preliminary…
  • The decision addresses memorization, Retrieval Augmented Generation (RAG), and the legal status of AI training.
  • The articles ANI cited were mostly from August and September 2024, so they couldn't have been part of the training data.

What happened

In an interim ruling, the Delhi High Court rejected a request by Indian news agency Asian News International (ANI) for a preliminary injunction against OpenAI. ANI, one of India's largest news agencies, sued OpenAI over its use of copyrighted material for AI training and in ChatGPT's outputs. Judge Amit Bansal denied the requested relief on both claims.

The court ran a three-part fairness test and sided with OpenAI on all three counts. OpenAI's use of ANI's works was limited to training, since no memorization or reproduction was proven. ANI also couldn't show economic harm because the two companies operate in different sectors. Even when users ask ChatGPT about ANI headlines, the model only returns topics and, at most, a few article titles. S. cases including Bartz v.

Anthropic and Kadrey v. Meta, where language model outputs were deemed transformative. He also pointed to the earlier Google Books ruling. The judge also found that trained language models improve access to information, support education, advance scientific research, help with software development, enable translation, and create tools for people with disabilities. The Delhi ruling joins a growing list of court decisions worldwide that have reached conflicting conclusions.

The Intercept, on the other hand, won a partial victory through a DMCA complaint over copyrighted material that had been stripped out before training. In Ross Intelligence v. Thomson Reuters, a court denied fair use because the AI research tool directly competed with Thomson Reuters' legal database Westlaw, making the use non-transformative.

The court stressed that this ruling applied only to this non-generative use case and couldn't be extended to large language models. In the Anthropic case, a federal court in San Francisco called AI training with copyrighted works "spectacularly" transformative, a strong signal favoring fair use. But the court drew a "Napster comparison" because Anthropic had used pirated books from shadow libraries as training data.

Why it matters

The decision addresses memorization, Retrieval Augmented Generation (RAG), and the legal status of AI training. AI copyright law expert Andres Guadamuz calls the ruling an important early win for OpenAI. ANI submitted several ChatGPT outputs to the court that it claimed were substantial copies of its articles. The move backfired because OpenAI showed that the models used, GPT-4 and GPT-4o, were trained on data from April 2022 and April 2024.

The articles ANI cited were mostly from August and September 2024, so they couldn't have been part of the training data. The judge's preliminary view was that the similarities came from RAG, which lets a language model retrieve online information in real time, much like a search engine. ANI hadn't addressed RAG in its filing, so the court couldn't make a final ruling on the issue.

The judge said RAG-based outputs could qualify as "communication to the public," a question the court will address in the main proceedings. ANI's case got weaker from there. " Even so, ANI couldn't produce a single verbatim copy. The judge found that facts in news articles in general aren't copyrightable and that reproducing topics and headlines didn't amount to direct competition with ANI in this case.

The evidence also didn't support ANI's claim that OpenAI permanently stores training data in its models and can reproduce the agency's work verbatim on demand. But the court will revisit that question in the main proceedings. ANI also failed to show that copying its work for AI training amounted to copyright infringement.

Both parties agreed that OpenAI had used ANI content during training, but OpenAI argued that the material made up a tiny share of the overall dataset and that the model extracted only non-expressive elements such as grammar, syntax, and language patterns. The judge looked at exceptions under Indian copyright law and relied on a clause covering "private or personal use, including research," reading "research" broadly enough to cover AI training.

For that exception to hold, the judge set conditions. Training copies must come from lawful sources, not shadow libraries or paywalled sites accessed without permission. OpenAI also never made the training copies public and processed them only internally. Guadamuz says this is the first time a court has explicitly found that AI training falls under a private use exception.

What to watch

Fair use doesn't cover unlawfully obtained material. 5 billion to book authors for using those pirated copies. S. Copyright Office rejected the AI industry's argument that training on "vast troves of copyrighted works" broadly qualifies as fair use. The official who wrote the report was fired by the Trump administration shortly after it was published. In Europe, two courts reached opposite conclusions almost simultaneously.

The Munich Regional Court ruled in the GEMA case that song lyrics were reproducible in the model weights, making it a copyright-relevant reproduction.