Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training examples verbatim and without authorship attribution, raising legal and ethical concerns around plagiarism and license compliance. Classical fingerprint-based plagiarism detectors based on fingerprinting, such as Winnowing, remain highly effective, yet the inspection requires comparing fragments of code to the entire training set, and their linear-time se
Record details
Published: 27 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Privacy · Copyright & IP
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Challenges to Grassroots Organization Engagement with AI Policy
arXiv · 18 June 2026
AI and Doctrinal Collapse
Harvard Berkman Klein Center · 30 June 2026
TILDE: TILt-based Distributional Erasure for Concept Unlearning
arXiv · 7 July 2026
Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition
arXiv · 9 April 2026
Central Bank of Kenya approves 25 new loan apps
Techpoint Africa · 15 July 2026
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
arXiv · 27 May 2026
How to cite this record
ethics.ai (27 May 2026), “Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets,” evidence record 3588, https://ethics.ai/record/3588 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.