Topic · updated daily · RSS feed for this topic
Copyright & IP
NYT v. OpenAI and every AI copyright fight: training data, fair use, licensing deals and artist rights, updated daily.
The New York Times is escalating its fight with OpenAI, urging a judge to impose sanctions
News outlets are asking a federal judge to sanction OpenAI, claiming it’s hiding ChatGPT logs relevant to a landmark copyright trial.
The New York Times is escalating its fight with OpenAI, urging a judge to impose sanctions
The New York Times , the Daily News and other media outlets are asking a federal judge to impose sanctions on OpenAI , escalating a fight over artificial intelligence and copyright that could shape the future of a struggling news industry . The newspapers allege the ChatGPT maker is hiding evidence important to what could be a landmark copyright infringement trial over how OpenAI and its business partner, Microsoft , built their AI technologies using millions of news articles. At issue is whethe
The New York Times is escalating its fight with OpenAI, urging a judge to impose sanctions
News outlets are asking a federal judge to sanction OpenAI, claiming it’s hiding ChatGPT logs relevant to a landmark copyright trial.
The New York Times is escalating its fight with OpenAI, urging a judge to impose sanctions
News outlets are asking a federal judge to sanction OpenAI, claiming it’s hiding ChatGPT logs relevant to a landmark copyright trial.
The New York Times is escalating its fight with OpenAI, urging a judge to impose sanctions
News outlets are asking a federal judge to sanction OpenAI, claiming it’s hiding ChatGPT logs relevant to a landmark copyright trial.
It’s Only Fair Use the First Time… Right? Debunking Copyright Urban Legends
This guest post, by Katherine Klosek of ARL and Stephen Wolfson of the University of Pennsylvania, is the latest in […]
Vietnam clarifies AI authorship, training data and copyright liability: A comparative lens
This article was originally published by IAPP linked here. Vietnam’s approach to artificial intelligence regulations crosses many topics and sectors, with a common theme emerging: human‑centered, state‑supervised and legally accountable. The country’s policy direction is clearly reflected in its first standalone Law on Artificial Intelligence, which took effect March 2026. At the same time, Vietnam [...] The post Vietnam clarifies AI authorship, training data and copyright liability: A comparati
OpenAI may have made a fatal misstep in copyright fight with news orgs
OpenAI may be sanctioned for hiding, deleting ChatGPT logs in NYT copyright fight.
[MàJ] Springer Nature a republié les articles rétractés de Max Planck
Des chercheurs québécois ont remarqué que deux articles du chercheur allemand Max Planck datant des années 1940 ont été rétractés par l’éditeur scientifique. La rétractation non datée, elle, accusant le physicien de violation de copyright, serait en fait due à une détection automatique zélée. [Mise à jour le 9 juillet à 9h30] : Un mois […]
Intellectual property strategies for AI-enabled health innovation
Navigating intellectual property (IP) issues for new innovations has never been simple. With healthcare solutions based on artificial intelligence (AI), it can be especially complex. Innovators, ...
Springer Nature un-retracts Planck papers, citing “human error”
Today the Retraction Watch list of Nobelists who have retracted papers bids Verabschiedung to Max Planck. After days of scrutiny, Springer Nature has restored two papers by Planck, who won the Nobel for Physics in 1918, reversing a 2011 decision to retract the articles for “copyright violations.” Both articles are back, and now carry the … Continue reading Springer Nature un-retracts Planck papers, citing “human error”
TILDE: TILt-based Distributional Erasure for Concept Unlearning
Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to suppress unwanted concepts after training. Existing methods often remove the target concept effectively, but practical unlearning also requires an equally fundamental property: the unlearned model should retain quality, diversity, and semantic coverage on benign generat
Pluralistic: CARDiac, syntax coloring, view source and vibe code (03 Jul 2026)
Today's links CARDiac, syntax coloring, view source and vibe code: With great abstraction comes great power comes great responsibility comes great loss of fidelity. Hey look at this: Delights to delectate. Object permanence: Real elections v reality TV; Copyright troll loses license; Who gets fed housing subsidies? Trump x forced labor. Upcoming appearances: London, Edinburgh, Sydney, Melbourne, Brighton, London, South Bend. Recent appearances: Where I've been. Latest books: You keep readin' em,
What's on in the Lords 29 June
On Tuesday, the Communications and Digital Committee continued looking into AI and copyright, speaking to Liz Kendall MP, Secretary of State for Science, Innovation and Technology and Lisa Nandy MP, ...
Priority Open Recommendations: Nuclear Regulatory Commission
What GAO Found In May 2025, GAO identified nine priority recommendations for the Nuclear Regulatory Commission (NRC). Since then, NRC has not implemented any of these recommendations. In May 2026, GAO identified two additional priority recommendations, bringing the total to 11. GAO is highlighting the following three areas that warrant timely and focused attention: Addressing the security of radiological sources, Improving risk-informed decision-making, and Licensing advanced nuclear reactors. A
AI and Doctrinal Collapse
Visiting Scholar Alicia Solow-Niederman identifies "inter-regime doctrinal collapse," a source of legal strain by which the boundaries between information privacy law and copyright law become ...
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law
Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards focus narrowly on verbatim memorisation. EU copyright doctrine applies a broader standards: substantial similarity, which extends to stylistic choices, narrative structure, and creative elaboration. This mismatch between what current methods detect and what the law protects leaves a significant compliance gap. We introduce PSALM, an LLM-as-a-judge framework that
The US now has a de facto model licensing system
OpenAI’s latest model, GPT-5.6, is waiting for government approval.
What Should Be Done
How to get past improvised model licensing
Challenges to Grassroots Organization Engagement with AI Policy
Public policies are being developed around the world to address privacy, economic, intellectual property, energy, and other risks that AI technologies pose. Involvement from the general public is essential to governance as an accountability and alignment mechanism. However, participating in and impacting policymaking can be challenging for sections of the public that lack extensive networks, lobbying capabilities, and other forms of power. This challenge is especially acute for marginalized comm
Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
Progress in legal AI increasingly depends on access to authoritative legal text at scale. Yet one of the most consequential layers of American law remains largely absent from existing machine-readable corpora: local ordinances. Local codes govern zoning, housing, business licensing, public health, noise, animal control, and many other domains of everyday regulation, but they are fragmented across vendor platforms designed for human browsing rather than bulk research access. We introduce LOCUS -
Digital Speech Acts Retain Control of Copyright with People, Not Platforms
Legal precedents protect computer code as copyrightable expression. They have enabled centralized digital platforms -- operating from corporate servers that hold all user data -- to construct private governance regimes through the interaction of copyright, contract, and technical architecture: people who create virtually all platform value must surrender effective copyright control through Terms of Service agreements as a condition of participation. In contrast, grassroots platforms consist of c
Market Design for AI: Beyond the Copyright Binary
How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves individual incentives for high-quality content creation? Existing approaches take polar positions: a "free-for-all" model based on fair use and a "strong intellectual property rights" model. We show that both fail: Free-for-all does not compensate creators, and -- by modeling as a static Stackelberg game -- strong intellectual property rights also underpower
Bypassing Copyright Protection in Diffusion-based Customization via Two-Stage Latent Feature Optimization
With the growing concerns over copyright infringement in diffusion-based customization, adversarial attacks have emerged as a prominent defense strategy to prevent malicious content forgery in personalized image generation. However, current defenses typically introduce persistent perturbations in the latent space of Latent Diffusion Models (LDMs), which remain susceptible to adaptive bypasses by adversaries. In this paper, we introduce Two-Stage Latent Feature Optimization (TS-LFO), an efficient
Auditing Training Data in Domain-adapted LLMs: LoRA-MINT
We present LoRA-MINT, a new methodology for Membership Inference Test (MINT) applied to recent Large Language Models (LLMs) fine-tuned for specific Natural Language Processing (NLP) tasks through Low-Rank Adaptation (LoRA). The primary goal is to assess whether individual samples were part of the training data of these adapted models, providing a useful auditing tool for the management of intellectual property and sensitive data. Our analysis explores the relationship between model perplexity an
The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol
Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from scratch. These decisions, whether to build functionality from scratch or buy into an external library, hereafter build-versus-buy, carry direct consequences for software security, licensing compliance, performance, and long-term maintainability. Yet no controlled experimental study has examined what governs build-versus-buy decisions in agentic AI c
Reproducibility is the New Copyleft: Defining AGI-oriented Reproducible Builds
Copyleft, as implemented in licenses such as the GNU General Public License, was a legal hack that used copyright to guarantee user freedom by tying the availability of source code to every act of distribution. Its normative force rested on an implicit technical premise: that source code and object code stand in a well-defined, humanly auditable, and reproducible relationship. Large language models and, prospectively, Artificial General Intelligence (AGI) systems systematically violate this prem
A Framework for Graph-Conditioned Hierarchical Shapley Attribution in Patent Valuation
Estimating the economic contribution of a single patent inside a product that embodies tens of thousands of patents is a long-standing unsolved problem in intellectual property economics. We propose PatentXAI, a framework that treats patent valuation as a problem of explainable AI: given a characteristic function v(S) encoding the revenue achievable by patent subset S, a patent's Shapley value measures its fair share of product profit in a way that satisfies efficiency, symmetry, dummy, and addi
Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training examples verbatim and without authorship attribution, raising legal and ethical concerns around plagiarism and license compliance. Classical fingerprint-based plagiarism detectors based on fingerprinting, such as Winnowing, remain highly effective, yet the inspection requires comparing fragments of code to the entire training set, and their linear-time se
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance requires auditable reasoning, safety and ethics alignment, and resilience to adversarial misuse. Here we present SafeMed-R1, trained with a traceable Clinical Trust Signals(CTS) pipeline that links each reasoning instance to clinician rubric scores and edit histories, and aligned through safety and ethics supervision and red team stress testing.
From Licensing to Open Access: Designing a Sustainable Transition in Operational Weather Data
This translational article documents the European Centre for Medium-Range Weather Forecasts (ECMWF) transition from a restricted data licensing model to open access under CC BY 4.0, completed in October 2025. The policy context included EU open data requirements and alignment with international data exchange frameworks. The transition was implemented through a tiered service model that kept core forecast data open while offering operationally supported delivery as a cost-recovered service. Betwe
CATA: Continual Machine Unlearning via Conflict-Averse Task Arithmetic
Vision-language models (VLMs) have shown remarkable ability in aligning visual and textual representations, enabling a wide range of multimodal applications. However, their large-scale training data inevitably raises concerns about privacy, copyright, and undesirable content, creating a strong need for machine unlearning. While existing studies mainly focus on single-shot unlearning, practical VLM deployment often involves sequential removal requests over time, giving rise to continual machine u
Jurisdiction over Ubiquitous Copyright Infringements: Should Right-Holders Be Allowed to Sue at Home?
The Internet, and more recently cloud computing, has transformed the technological, economic, social, and cultural conditions under which intellectual property rights are exploited. These developments also challenge traditional rules of private international law, particularly rules governing international jurisdiction. This paper examines when courts should assert jurisdiction over cross-border copyright disputes arising in cloud-based environments. It focuses on the risks faced by right holders
Position: The Term "Machine Unlearning" Is Overused in LLMs
Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/licensing disputes, and safety or product-policy requirements. This position paper argues that machine unlearning is overused as a term in LLM research and should be reserved for dataset-defined deletion: removing the training influence of a precisely specified forget set such that the resulting model is approximately indistinguishable from retraini
In-Context Credit Assignment via the Core
We propose incentive-aligned mechanisms for in-context credit assignment: the task of assigning credit for AI-generated content (e.g. code, news articles, short-form videos) among creators whose intellectual property appears in the context window. Our approach is based on the least core solution concept from cooperative game theory, which distributes value in a way that is as stable as possible by ensuring that no subset of creators is significantly under-compensated relative to the value they c
Can Code Evaluation Metrics Detect Code Plagiarism?
Source Code Plagiarism Detection (SCPD) plays an important role in maintaining fairness and academic integrity in software engineering education. Code Evaluation Metrics (CEMs) are developed for assessing code generation tasks. However, it remains unclear whether such metrics can reliably detect plagiarism across different levels of modification (L1-L6), increasing in complexity. In this paper, we perform a comparative empirical study using two open-source labelled datasets, ConPlag (raw and tem
ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encoders
Self-supervised learning (SSL) encoders are invaluable intellectual property (IP). However, no existing SSL watermarking for IP protection can concurrently satisfy the following two practical requirements: (1) provide ownership verification capability under black-box suspect model access once the stolen encoders are used in downstream tasks; (2) be robust under adversarial watermark detection or removal, because the watermark samples form a distinguishable out-of-distribution (OOD) cluster. We p
Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition
High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory and copyright constraints. This scarcity hampers model development--ironically, in settings where generative models are most needed to compensate for the lack of data. This creates a self-reinforcing challenge: limited data leads to poor generative models, which in turn fail to mitigate data scarcity. To break this cycle, we propose a reinforcement
Agentic Copyright, Data Scraping & AI Governance: Toward a Coasean Bargain in the Era of Artificial Intelligence
This paper examines how the rapid deployment of multi-agentic AI systems is reshaping the foundations of copyright law and creative markets. It argues that existing copyright frameworks are ill-equipped to govern AI agent-mediated interactions that occur at scale, speed, and with limited human oversight. The paper introduces the concept of agentic copyright, a model in which AI agents act on behalf of creators and users to negotiate access, attribution, and compensation for copyrighted works. Wh
Can AI be a moral victim? The role of moral patiency and ownership perceptions in ethical judgments of using AI-generated content
The growing use of generative AI raises ethical concerns about authorship and plagiarism. This study examines how people judge the reuse of AI-generated content, focusing on moral patiency and ownership perceptions. In an experiment, participants evaluated two substantively similar manuscripts in which the original source was described as authored by a human, an AI system, or an AI agent with a human-like name. Results showed that copying AI-generated work was judged less unethical, less plagiar