Copyright and artificial intelligence before the CJEU: Case C-250/25 (Like Company v. Google)
The Court of Justice of the EU is called for the first time to rule on the application of copyright to the training and operation of generative AI models, in Case C-250/25 (Like Company v. Google Ireland). Analysis of the four questions for a preliminary ruling and the operational implications for data sourcing, licensing, and M&A due diligence, in light of Law 132/2025 and the AI Act.
— Studio LX20 Law Firm
Case C-250/25, _Like Company v. Google Ireland_ — the first reference for a preliminary ruling in the Union on generative AI and copyright. Operational implications for data sourcing strategies, licensing contracts, and M&A due diligence.
The case, in summary
On 3 April 2025, the Budapest-Capital Regional Court (Budapest Környéki Törvényszék, not to be confused with the Metropolitan Court) referred to the Court of Justice of the European Union the first request for a preliminary ruling dedicated to the intersection of generative AI and copyright. The dispute is between a Hungarian publisher, Like Company, and Google Ireland and concerns the Gemini chatbot (formerly Bard): the publisher complains that the model's responses reproduce portions of its press articles beyond the limit of a "very short extract" and that the training of the model itself involved unauthorized reproductions of its content.
The Court will not decide on the merits of the Hungarian dispute: it will answer four interpretative questions intended to apply throughout the single market. The first oral hearing was held on 10 March 2026; the Advocate General's opinion is expected in autumn 2026; the judgment is not expected before the end of 2026, more likely in 2027.
A widespread misconception must be corrected immediately: there is no "Turnitin" case nor any proceedings with a different case number in this matter. The reference to be cited is only one: C-250/25.
The four questions
The questions for a preliminary ruling touch upon the two pillars of EU digital copyright law — Directive 2001/29/EC (InfoSoc) and Directive (EU) 2019/790 (DSM) — along the entire value chain of a language model: input, training, output.
- Communication to the public. Whether the response of a chatbot that reproduces a text partially identical to a press article, to an extent exceeding the use of "individual words or very short extracts", constitutes an act of communication to the public relevant under Art. 15 DSM and Art. 3(2) InfoSoc.
- Training as reproduction. Whether the training process — tokenization and learning of linguistic patterns from protected works — constitutes "reproduction" within the meaning of Art. 2 InfoSoc.
- Text and data mining exception. If so, whether such reproduction falls under the TDM exception of Art. 4 DSM (lawful access and absence of a reservation by the rightholder).
- Output as an act of the provider. Whether the generation of a response that reproduces, in whole or in part, a protected content in response to a user's prompt constitutes a reproduction attributable to the service provider.
What is at stake is the legal qualification of the entire pipeline. Google argues that Gemini does not store or retrieve copies, but generates text probabilistically, and that any similarity is incidental or the result of "hallucination", invoking the exceptions for temporary copies (Art. 5(1) InfoSoc) and for TDM (Art. 4 DSM). The publisher replies that both training and operation exceed the limits of the law.
A point that emerged from the hearing deserves attention because it goes beyond the individual case: the issue of the territorial application of EU copyright law. Several Member States (Hungary, Denmark, Greece, Spain, France) have argued for a "unitary" reading of the process — training, grounding, input, output, and communication as a single act — such as to bring EU rules into play even when training takes place outside the Union, as long as the system is marketed in the internal market. Germany has taken the opposite position, anchoring territorial relevance to a concrete act of reproduction on a medium within a Member State. The Commission, for its part, has suggested that the questions may be inadmissible, deeming them abstract. It is therefore prudent not to take the outcome on the merits for granted: the Court might not answer everything.
Why it matters for Italian law
On a substantive level, domestic law coincides with Union law, because the transposition has been faithful. Legislative Decree No. 177 of 8 November 2021 (in force since 12 December 2021) introduced two provisions into Law No. 633 of 22 April 1941 (LDA):
- Art. 70-ter, which transposes the scientific research exception of Art. 3 DSM, reserved for research organisations and cultural heritage institutions, non-derogable and therefore not subject to opt-out;
- Art. 70-quater, which transposes the general exception of Art. 4 DSM: TDM on lawfully accessible materials is permitted by default, unless reserved (opt-out) by the rightholder.
To this architecture was added, in 2025, an intervention that makes it central for those who train or use AI in Italy. Law No. 132 of 23 September 2025 (in force since 10 October 2025):
- with Art. 25, it amended Art. 1 LDA, clarifying that works created with the aid of AI tools remain protected as long as they are the result of a real human intellectual contribution (the so-called "humanity reserve"), and introduced Art. 70-septies, which expressly anchors reproductions and extractions of text and data carried out by means of AI models and systems — including generative ones — to the rules of Arts. 70-ter and 70-quater;
- with Art. 26, it had an impact on the criminal level, inserting letter a-ter) into Art. 171 LDA, which sanctions those who reproduce or extract text or data in violation of Arts. 70-ter and 70-quater, including by means of AI systems.
The Italian legislator has thus expressly resolved an interpretative doubt (the applicability of the TDM exceptions to AI training) which at the Union level is precisely one of the questions submitted to the Court in C-250/25. However, the issue of enforcement remains open, in Italy as elsewhere: the opt-out is a right, the proof of whose violation is held by the party against whom it would be asserted, i.e., the one who trains the model.
The opt-out and its "appropriate manner"
The effectiveness of the reservation under Art. 70-quater depends on its form. For content made available online, the recitals of the DSM clarify that the reservation is appropriate only if expressed by machine-readable means: metadata, exclusion protocols (typically the robots.txt file), site terms and conditions in machine-readable format, digital rights management systems. In other cases, contractual instruments or a unilateral declaration are adequate.
This level is combined with the AI Act (Regulation EU 2024/1689), Art. 53, which requires providers of general-purpose AI models to put in place a policy to respect Union copyright law, including respecting reservations under Art. 4(3) DSM. In practical terms: the AI Act transforms the opt-out from a mere private-law option into a compliance obligation for the model provider.
Operational implications
For those who develop or train models (data sourcing). The combination of the expected outcome of C-250/25, Art. 70-septies LDA, and Art. 53 AI Act shifts the center of gravity from the logic of "everything online is usable" to a logic of documented provenance. It is advisable, from now on, to: map datasets distinguishing between sources with lawful access and uncertain sources; implement an automated mechanism for detecting and respecting opt-outs (including textual reservations and robots.txt files); keep evidence of the supply chain (data provenance), because in litigation or regulatory audits the burden of proving lawfulness tends to fall on the one who trains.
For those who license or hold content (rightholders and publishers). For the reservation to be enforceable, it must be exercised in a technically appropriate and documented form. Contractual clauses prohibiting TDM must be coordinated with machine-readable measures: a reservation contained only in a contract may not be enforceable against those who perform scraping without a contractual relationship. On the active side, the possible qualification of the output as a reproduction attributable to the provider (question 4) would strengthen the negotiating position of those who license content for training.
For M&A and investments in AI-driven targets. Due diligence on a company that develops or integrates models must include a specific verification of the rights to the training data: origin of the datasets, existence of licenses, management of opt-outs, exposure to ongoing or prospective copyright litigation. A restrictive outcome from the Court may affect the valuation (need to re-license datasets, compliance costs, risk of injunctions) and suggests the inclusion of targeted representations & warranties and, where appropriate, indemnification or price adjustment mechanisms.
The picture is not limited to the CJEU
The Hungarian reference is part of a moving European jurisprudential context. In Germany, the Munich I Regional Court (judgment of 11 November 2025, GEMA v. OpenAI) held that the storage of song lyrics in GPT models constitutes unlawful reproduction, rejecting the defense based on TDM; the earlier Kneschke v. LAION case had already addressed the research exception. As these are transposing rules homologous to the Italian ones, these trends are also significant for the national interpreter. The direction, although not consolidated, is towards a restrictive interpretation of the exceptions — consistent with CJEU case law (Infopaq, Pelham) on the strict interpretation of limitations to copyright.
Operational conclusions, and the limits to keep in mind
At present, three recommendations withstand critical scrutiny and do not depend on the outcome of the specific case:
- Copyright compliance for training can no longer be postponed pending the judgment. Art. 53 of the AI Act and Law 132/2025 (Arts. 70-septies and 171, letter a-ter, LDA) are already existing law in Italy, with criminal aspects as well. C-250/25 will clarify the Union framework, but the obligation to have policies and respect opt-outs exists today.
- Documented data provenance is the main defensive safeguard. Regardless of how the Court qualifies training, those who can demonstrate lawful access and respect for reservations substantially reduce their exposure.
- In M&A transactions, rights to training data must be treated as a separate asset to be verified, not as a technical detail.
Three caveats, for the sake of analytical honesty. First: the outcome on the merits is uncertain. The Commission has raised the inadmissibility of the questions and the Court might only answer partially, leaving the most relevant points for the industry open. Second: the territorial issue is still undefined; a "unitary" reading would extend the effectiveness of EU rules even to extra-EU training for systems marketed in the internal market, but Member States are divided and the solution cannot be predicted with certainty. Third: the structural evidentiary deficit of the opt-out persists — a right whose violation is difficult for the rightholder to prove. Anyone who bases a strategy on the assumption of a clear-cut outcome, one way or the other, takes a risk that cannot be quantified with certainty today. The prudent recommendation is to build processes that are robust to both scenarios.
This contribution is for informational purposes only and does not constitute legal advice. The legislative and case law sources are updated as of June 2026. Laws and cases cited: Directive 2001/29/EC (Arts. 2, 3, 5); Directive (EU) 2019/790 (Arts. 3, 4, 15); Legislative Decree 177/2021 (Arts. 70-ter and 70-quater of Law 633/1941); Law 132/2025 (Arts. 25 and 26; Arts. 1, 70-septies and 171 of Law 633/1941); Regulation (EU) 2024/1689 (Art. 53); CJEU Case C-250/25, Like Company v. Google Ireland.