top of page

Naver AI Copyright Lawsuit Turns on What Its News Contracts Actually Allowed

Sep 9
12 min read

Naver defended its AI training practices in court despite three broadcasters claiming their articles were used without permission. The Naver AI copyright lawsuit now centers on whether older news-sharing contracts authorized uses that publishers say they never knowingly approved.

KBS, MBC, and SBS sued Naver in January 2025 over material allegedly used to train HyperCLOVA and HyperCLOVA X. They seek damages and an order stopping further training with their news content.

Naver told the Seoul Central District Court that contractual language covering research, service development, and sentence extraction permitted its use. According to September 8 coverage, the company maintained that these agreements protect it from the broadcasters' claims.

That defense creates a sharper dispute than the familiar argument over whether machine learning qualifies as fair use. The court must also decide what the parties licensed before generative AI became a defined commercial category.

The broadcasters say authorization for distributing news through Naver did not amount to permission for building a commercial language model. Naver argues that the contracts anticipated broader technical processing, even if they did not name large language models.

The distinction matters beyond one Korean lawsuit. Technology companies often hold years of licensed material under agreements written before current AI systems existed. If broad development clauses cover model training, those contracts become valuable defenses. If informed consent was required, companies face new licensing costs and possible restrictions.

Naver Says Its Contracts Covered AI Training

The immediate change is not a new lawsuit, but a more developed contractual defense from Naver after months of court scrutiny.

MLex reported on September 8, 2026, that Naver again told the South Korean court its use complied with content-sharing contracts. The hearing followed a significant August proceeding focused on how those agreements were negotiated.

The litigation began when KBS, MBC, and SBS accused Naver of using their news articles to train HyperCLOVA and HyperCLOVA X. The broadcasters brought copyright, unfair competition, and related claims in the Seoul Central District Court.

The broadcasters originally sought 200 million won each, or 600 million won in total, according to Korean reporting. Their requested relief also includes an injunction against training on the disputed material.

Those amounts do not capture the case's full commercial importance. An injunction or restrictive interpretation could affect data already incorporated into Naver's model-development process.

The broadcasters represent South Korea's three major terrestrial television networks. The Korean Broadcasters Association, which represents 39 broadcasters, supported the action and had previously demanded separate AI compensation discussions.

According to the initial lawsuit filing, the association notified Naver, Kakao, and Google Korea about its position in December 2023. It said news content required separate permission and compensation when used for AI learning.

The association's task force later requested information about Naver's sources, training material, and acquisition methods. It said Naver did not provide a clear disclosure.

Naver initially said it needed to review the complaint before presenting a detailed position. Its courtroom defense has since become more specific.

Naver points to provisions in its agreements concerning research, new service development, and technical processing. It also cites contractual references to an AI platform and sentence extraction.

Sentence extraction can describe taking individual passages from larger documents for indexing, analysis, or another computational process. Naver argues that the term reflected data processing associated with machine learning.

The broadcasters reject that interpretation. They say extracting text for an existing news service differs from copying articles into a model-training corpus.

That difference creates the central contractual question. Permission to process content does not automatically answer which processing purposes the parties accepted.

Naver also argues that certain factual news reports do not qualify for full copyright protection under Korean law. It has challenged the broadcasters to identify the protected expression and specific works underlying their claims.

The broadcasters answer that their articles have economic value and contain protectable selection, structure, and expression. They characterize Naver's alleged copying as systematic commercial use, not incidental reference to isolated facts.

No final ruling has established which interpretation governs. The September hearing continued an active dispute rather than resolving Naver's liability.

The Naver AI Copyright Lawsuit Pressures Both Sides

The broadcasters need to show that ordinary news distribution rights stop before model training, while Naver must justify a broader reading of older contracts.

For Naver, the immediate pressure falls on the legal foundation beneath its Korean-language AI program. HyperCLOVA X depends on large collections of Korean, English, and code material.

Naver has not publicly identified the disputed articles within its training corpus. Its published model documentation describes data categories at a higher level.

The company's 2024 technical report says HyperCLOVA X used a balanced mixture of Korean, English, and code data. It does not publicly enumerate every publication represented in that mixture.

That disclosure gap gives the broadcasters a practical problem. They must identify infringement without complete access to Naver's training records.

They reportedly presented answers generated by HyperCLOVA X as evidence that broadcaster content had been used. Naver disputed that method because language models can produce unreliable statements about their own training.

A model's answer about its data sources is not a dependable audit record. Large language models generate likely text rather than retrieving a verified inventory of their training documents.

The court therefore faces a discovery challenge alongside the legal dispute. It must determine how plaintiffs can identify copied works when the defendant controls the relevant data records.

Naver has argued that the broadcasters did not adequately specify which works were infringed. The broadcasters say the unusual scale and opacity of AI training make greater specificity difficult.

This is more than a procedural disagreement. The party controlling training records can hold a decisive evidentiary advantage before a court orders disclosure.

The broadcasters face pressure of their own. They distributed content through Naver for years and accepted agreements allowing technical uses beyond simple display.

A court might conclude that broad development language covered machine learning as it existed when the contracts were signed. The broadcasters would then need another legal theory to separate generative training from licensed platform development.

Their position also requires distinctions among different kinds of news. Korean copyright law excludes certain current-events reports consisting of simple factual communication from protection.

That exclusion does not mean every news article is unprotected. Reporting can include original expression, analysis, organization, photography, and other creative elements.

Naver has pressed this distinction by asking the plaintiffs to identify protectable material. That approach can narrow a mass claim into article-by-article questions.

The broadcasters instead emphasize the commercial value of the collected corpus. Their claim treats the dataset as a resource assembled through newsroom investment, not merely a list of public facts.

The result pressures both sides toward evidence they have not fully disclosed publicly. Naver needs to connect contract language to the specific training activity. The broadcasters need to connect protected works to the models.

Old Contracts Meet a New Kind of Commercial Use

The case turns on whether broad technical permission can stretch from improving a news platform to training a reusable generative model.

The relevant agreements did not emerge in the ChatGPT era. Testimony examined provisions and special terms negotiated between 2017 and 2020.

Naver says references to Clova, AI platforms, research, and sentence extraction placed the broadcasters on notice. It argues that machine learning was already familiar when the parties negotiated those terms.

A Naver employee involved with news partnerships between 2016 and 2023 testified during the sixth hearing. The employee said discussions occurred through email, meetings, and telephone calls.

The witness maintained that Naver had adequately explained the special nature of its AI platform. The witness also said the broadcasters' legal teams closely reviewed contractual language and raised questions when terms were unclear.

The broadcasters disputed whether those discussions covered generative model training. They emphasized that terms such as large language model and generative AI did not appear in the agreements.

The August hearing record describes a 2020 clause allowing Naver to use information for research, service improvement, and new service development. Media organizations objected to its breadth, and Naver later added a prior-consent requirement.

That revision complicates both narratives. Naver can point to an established contractual process surrounding technical use. The broadcasters can argue that the later consent requirement shows unrestricted use was unacceptable.

Naver also reportedly argued that broader rights produced higher payments to the broadcasters. During an earlier hearing, it said payments increased by 50 percent and reached hundreds of millions of won over five years.

Those figures remain part of Naver's courtroom position, not a judicial finding about what the payments purchased. A larger payment does not establish AI consent unless the contract or negotiating record links the two.

The broadcasters focus on technological and commercial context. They say pre-2020 discussions occurred before ChatGPT made general-purpose generative systems a recognizable market.

Machine learning was certainly established by then. However, classifying articles or improving search differs economically from training a model that can generate new text across many services.

This is the lawsuit's core tradeoff. Contract language must remain flexible enough to cover normal technical development, yet specific enough to preserve meaningful limits.

An agreement listing every future computational method would quickly become obsolete. An unlimited development clause could also transfer unforeseen rights without a publisher understanding the exchange.

Courts often interpret contracts using their text, purpose, negotiating history, and the parties' conduct. The Naver dispute concentrates all four inquiries around a technology whose market role changed rapidly.

Naver links the contracts to Clova and established AI work. The broadcasters link them to news distribution and service improvement, not general-purpose model development.

The court's August comments increased uncertainty for Naver. After hearing testimony, the judges indicated that generative AI training had probably not been specifically explained to the broadcasters.

That observation was not a final ruling. It still weakens any simple claim that the contracts clearly settled the issue.

The September defense appears designed to keep the written agreements at the center. Naver maintains that contractual authorization can exist without a detailed explanation of every future model architecture.

The broadcasters need the court to adopt a narrower boundary. Their case depends on model training being distinct enough to require separate, informed permission.

Copyright and Contract Defenses Do Different Work

Even if Naver loses the contract argument, it can still contest copyright protection, copying, fair use, and unfair competition.

The public debate often compresses AI copyright cases into one question: Is training fair use? This lawsuit contains several questions that operate independently.

First, the court must determine whether the broadcasters own copyright in the works they identify. Pure facts and basic current-events reporting receive different treatment from original expressive reporting.

Second, the broadcasters must establish that protected material entered Naver's training process. That issue depends on evidence about datasets, preprocessing, retention, and model development.

Third, the court must interpret the content-sharing contracts. If those agreements authorized the relevant use, copyright infringement claims can fail without a broad ruling on AI training.

Fourth, the court can examine statutory limitations such as fair use. South Korea's fair-use provision considers a use's purpose, the work's nature, the amount used, and market effects.

Those factors prevent an automatic answer. Training may transform text into statistical model parameters, but it also requires copying material during data preparation and computation.

Commercial purpose does not automatically defeat fair use. It can weigh against a defendant when combined with extensive copying and harm to a licensing market.

The broadcasters will likely emphasize substitution and lost licensing opportunities. They say AI developers should negotiate separate compensation for access to newsroom output.

Naver can argue that training produces a general language model rather than a database serving copies of articles. It can also distinguish functional processing from publishing the source works.

Output behavior remains relevant. A system that reproduces protected passages creates a different risk from one that uses material without returning recognizable expression.

The public record does not establish how frequently HyperCLOVA X reproduces broadcaster articles. It also does not reveal the quantity of material allegedly included.

That uncertainty limits strong conclusions about market harm. It also explains why contract interpretation has become such an attractive defense for Naver.

A contract can provide a cleaner path than defining fair use for an entire technology. The court could rule narrowly based on these parties' agreements.

However, even a narrow contractual ruling would influence other negotiations. Platforms and publishers would inspect similar clauses for references to research, extraction, service development, and machine learning.

Naver itself acknowledges that the litigation carries meaningful risk. Its 2026 offering circular identifies the broadcasters' action and says its outcome remains uncertain.

The filing says the plaintiffs allege unauthorized use of news content for generative AI training. It records their requests for monetary damages and injunctive relief.

Naver also warns investors that adverse intellectual-property decisions can produce damages, licensing obligations, changed business practices, or discontinued services. That language is a general risk disclosure, not a prediction about this case.

Still, it contradicts any impression that the contractual defense has already eliminated exposure. Naver has presented a defense, and the broadcasters are actively challenging it.

Publishers Are Testing Licensing Power, Not Just Ownership

The larger contest concerns who can set terms for high-value training data after years of platform distribution.

KBS, MBC, and SBS are both content owners and long-standing participants in Naver's news platform. That relationship separates this dispute from cases involving material gathered from the open web.

The broadcasters did not simply discover their articles on an unrelated service. They had contracts governing delivery, display, technical handling, and compensation.

Those agreements gave Naver lawful access to content. The unresolved question is whether lawful access for one purpose became authorization for another.

This pattern will recur wherever technology companies have deep content partnerships. Search engines, aggregators, cloud providers, and enterprise platforms often possess licensed archives collected under older terms.

Generative AI raises the value of those archives. Consistent, professionally edited language can support training, retrieval systems, evaluation, and product development.

Publishers now want to separate those uses from ordinary distribution. Their bargaining goal is not limited to damages for past copying.

A ruling requiring specific AI permission would strengthen demands for new licenses. It would also encourage detailed clauses covering training, fine-tuning, retrieval, evaluation, retention, and generated outputs.

A ruling favoring Naver would give platforms another path. Companies could rely on broad research or service-development clauses when those terms and surrounding negotiations support their interpretation.

Neither outcome would create a universal rule. Contract wording varies, as do payment arrangements and communications between partners.

The case also intersects with South Korea's ongoing effort to clarify AI copyright policy. Government work has considered text-and-data-mining rules, training-data disclosure, and systems for identifying rights information.

In February 2026, the Ministry of Culture, Sports and Tourism released guidance addressing fair use in generative AI training. Guidance can organize analysis, but courts still determine how the statute applies to contested facts.

The Korean dispute also has international parallels. News organizations have sued OpenAI and Microsoft in the United States, while other publishers have signed licensing agreements with AI companies.

Those strategies pursue the same scarce asset through different routes. Litigation tries to establish control after alleged use, while licensing converts access into a negotiated commercial relationship.

Naver's position is unusual because it says the negotiation already happened. The broadcasters answer that the bargain never included the use now at issue.

That disagreement makes the contracting process as important as the final wording. Evidence about presentations, emails, revisions, questions, and payments can show what each side reasonably understood.

The court previously declined to call a senior Naver executive requested by the broadcasters. It instead heard from an employee involved in news partnership work.

At the August hearing, the court indicated the broadcasters could proceed on the assumption that generative model training had not been specifically explained. It left open additional testimony from employees involved with the 2020 SBS agreement.

That procedural choice keeps the focus narrow. The judges appear interested in what the contracting teams communicated, not broad public statements about AI.

Naver says the broadcasters followed technology closely and understood machine learning's growing importance after AlphaGo's 2016 match. The broadcasters say general awareness of AI cannot substitute for permission.

The difference is fundamental. Knowing that a partner develops AI does not necessarily mean licensing content for training.

The reverse is also true. A contract does not become ineffective merely because a later technology commercializes an anticipated technical use more successfully.

The court must place the facts between those poles. That task makes this a contract case with copyright consequences, not simply another referendum on AI.

Three Signals Will Show Which Argument Is Winning

The next phase should reveal whether the court prioritizes contractual text, informed consent, or evidence from Naver's actual training pipeline.

The first signal is additional testimony or documentary evidence from the 2020 contract negotiations. The court previously raised the possibility of hearing from two SBS employees involved in that process.

Emails or drafts explicitly connecting research clauses with model training would strengthen Naver's defense. Records showing no such discussion would support the broadcasters' narrower interpretation.

The second signal is whether the court orders or receives more detailed training-data evidence. The plaintiffs need a credible link between their protected works and specific Naver models.

A dataset inventory, preprocessing record, or internal use description would carry more weight than a model's answer about its own training. Absence of that evidence would preserve Naver's challenge to the broadcasters' proof.

The third signal is how the judges sequence contract interpretation and fair use. A contract-based decision could resolve much of the dispute without establishing a general AI-training rule.

A broader fair-use analysis would matter to companies without content-sharing agreements. It would also influence South Korean publishers considering claims against other technology providers.

Readers should resist treating Naver's September position as a victory. The company says its contracts provide authorization, but the court has not accepted that conclusion.

The August proceeding exposed a weakness in Naver's narrative. The judges appeared unconvinced that generative AI training had been specifically explained when the agreements were made.

That weakness does not automatically invalidate broad contractual permission. It does make negotiating history and the exact language more important.

The Naver AI copyright lawsuit will become especially consequential if the court separates access from authorization. A company can lawfully receive content yet exceed the permitted purpose for using it.

Developers and enterprise buyers should watch that distinction in their own data agreements. A general right to analyze or improve a service may not clearly cover foundation-model training.

Publishers should examine the same clauses before assuming they retained every AI-related right. Existing definitions, technical schedules, amendments, and payment terms can change the analysis.

The practical response is careful recordkeeping. Teams need traceable datasets, documented permissions, defined uses, and a process for honoring later restrictions.

For anyone evaluating AI systems, the final question is concrete: Can the provider show where its training rights came from and what those rights actually allowed? Naver says its answer lies in contracts signed years ago. The broadcasters say those agreements never authorized the commercial use now before the court. The documents, negotiation record, and training evidence will decide which account holds.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page