HarperCollins Says Publishing Needs a Better Way to Prove Human Authorship
By Wendy Keir | EmpowerAi™
Something interesting is beginning to happen in publishing.
For the past couple of years, much of the discussion about AI-written books has focused on detection. A manuscript appears, someone suspects AI has been involved, a detector is run over the text and everyone then tries to decide what the percentage means.
That approach is beginning to look increasingly inadequate.
Brian Murray, the CEO of HarperCollins, has now said that publishers, agents and authors may need to develop an industry-wide approach to establishing whether a book has been written by a person. Speaking to Publishers Weekly following several high-profile disputes over suspected AI use, Murray said he does not believe there is a technical solution to the problem and suggested that authors may eventually need to retain records showing how their books were created.
I think that matters because it moves the conversation into a different place.
Instead of trying to prove authorship after something has gone wrong, publishing may need to start thinking about provenance as part of the normal process of creating a book.
AI detection is proving to be a fairly fragile foundation
There have now been several cases where books have become the subject of accusations about AI use.
The difficulty is that the technology used to investigate those accusations does not actually observe the writing process.
An AI detector examines patterns in a finished piece of text and estimates whether those patterns resemble material generated by a language model. It cannot see the research that happened before the manuscript was written. It cannot see the earlier drafts, the editing decisions or the years of thinking that may sit behind a particular paragraph.
Murray described the present situation as very unclear and said he does not believe detection software provides the answer. His view is that publishers, agents and authors will need to work together to develop better practices around establishing authorship before a book reaches publication.
That seems a much more useful direction.
If the only evidence available is the finished manuscript, everyone is trying to reconstruct a creative process that has already happened.
If evidence is retained while the book is being created, the picture becomes much clearer.
Authors may need to keep the history of the book
Most authors already leave a considerable amount of evidence behind when they write.
There are outlines, research documents, abandoned chapters, notes, emails with editors and earlier versions of the manuscript. There may be recordings of interviews, handwritten notebooks, client stories or pieces of an argument that appeared years earlier in articles and talks.
Until now, very few authors have thought of these things as evidence of authorship.
They were simply part of writing the book.
That may change.
Murray suggested that authors could be expected to retain records of their writing process, and I can see why publishers are beginning to think along those lines.
A version history showing how a chapter developed is much more informative than an AI detector analysing the finished prose.
Research notes can show where an argument came from. Previous writing can show that the author's ideas existed long before the book was completed. Editorial exchanges reveal the decisions that shaped the manuscript.
None of this provides perfect proof, and I doubt publishing will ever find a single test that does.
It does create a much more credible trail.
We also need to be careful about what we mean by AI-assisted
There is another part of this conversation that needs a little more precision.
Murray raised concerns that books written with AI might face copyright problems and suggested that an AI-assisted book could potentially be treated as though it were in the public domain.
The position set out by the US Copyright Office is more specific
Its guidance says that using AI to assist with creating a work does not automatically remove copyright protection. Human-authored expression can still be protected even where AI-generated material is also present. What cannot generally be protected is material produced entirely by AI where there is insufficient human creative control over the expressive elements.
That distinction matters enormously.
An author using ChatGPT to question the structure of a chapter is doing something very different from asking it to generate the chapter.
Using AI to identify repetition is different from asking it to create the underlying argument.
Using it to organise research is different from allowing it to invent the expertise.
We need better language around these differences because the phrase AI-assisted book can cover almost anything.
Expert authors have more provenance than they may realise
For expert authors, there is an interesting opportunity within all of this.
Their books usually do not begin with the manuscript.
The ideas may have developed over ten or twenty years.
A framework may have started with a client problem. It may have been refined through workshops, teaching, speaking and conversations. Some parts may have appeared in newsletters or courses long before they appeared in a chapter.
That accumulated history is intellectual provenance.
If publishing begins placing greater importance on establishing where ideas came from, expert authors are in quite a strong position.
They can often show the route by which the thinking developed.
The book becomes one expression of a larger body of intellectual property rather than an isolated collection of words.
That matters because generative AI makes polished writing increasingly easy to produce.
Someone can ask a model to create a credible-sounding framework in seconds. It can have three stages, a memorable acronym and a convincing explanation.
What it does not have is the history behind the framework.
There is no trail of clients who helped shape it. There are no years of testing what worked and what failed. There is no professional judgement that gradually altered the method.
The words may look similar.
The provenance is completely different.
AI Book Companion™ makes this distinction particularly important
This is one of the reasons I have become increasingly interested in the architecture behind an AI Book Companion™.
A Companion generates language every time a reader talks to it.
The actual sentence the reader receives may never have appeared in the book.
That means the system needs a very clear relationship with the author's intellectual property.
The book, frameworks, terminology, boundaries and approved supporting material form the source. The AI then uses those materials to help the reader understand or apply the author's thinking.
That distinction needs to remain visible.
The reader should know that they are talking to an AI system. They should also know whose body of work governs the conversation.
If the Companion explains a framework in another way, it needs to remain faithful to the author's meaning. If the source material does not contain an answer, the system should be able to acknowledge that rather than quietly importing something from elsewhere.
This is where provenance becomes part of the design.
You need to know which material belongs to the author, which sources were approved and where the Companion's authority ends.
The generated conversation sits on top of that foundation.
Provenance could become part of publishing infrastructure
At the moment, most of this happens informally.
Authors keep drafts because they are useful. Editors track changes because they need to manage revisions. Publishers keep correspondence because that is simply part of running a publishing business.
I can imagine some of that becoming more deliberate.
An author might keep a simple archive containing major manuscript versions, research notes and records of how AI was used.
Publishers might develop clearer declarations describing acceptable and unacceptable forms of AI assistance.
Agents might begin discussing AI use before submitting a manuscript rather than discovering the issue during due diligence.
None of this needs to become an enormous bureaucratic exercise.
The useful part is establishing a credible history of the work.
It also protects authors.
When an accusation is made, the author does not have to rely entirely on saying that they wrote the book. They have a record showing how it came together.
Given how unreliable AI detection can be, that may become increasingly important.
The interesting shift is from detection to evidence
Publishing has spent a lot of time asking how we detect AI-generated books.
There may never be a reliable answer to that question.
The models will continue improving, human writers will continue using AI in many different ways, and the boundary between assistance and generation will remain messy.
The more practical question may be how an author establishes the provenance of their work.
That can be answered.
They can retain drafts. They can document sources. They can preserve the history of their frameworks. They can be clear about how AI was used and which decisions remained theirs.
For expert authors in particular, there is already a substantial body of evidence sitting behind most good books.
Publishing may simply be reaching the point where that evidence becomes more important.
Three Key Insights
-
Major publishers are beginning to look beyond AI detection.
HarperCollins CEO Brian Murray has called for publishers, agents and authors to develop shared approaches to establishing human authorship, including potentially retaining records of the writing process.
-
AI assistance does not automatically remove copyright protection.
The US Copyright Office says human-authored elements can still receive copyright protection where AI has been used as a tool. Purely AI-generated material without sufficient human creative control is treated differently.
-
Provenance is becoming part of the value of expert intellectual property.
The history behind an author's frameworks, ideas and methodology creates evidence that generated prose alone cannot provide.
Three Questions I'm Thinking About
-
Will publishers eventually ask authors to provide some form of provenance record alongside a finished manuscript?
-
Do we need a much clearer vocabulary for distinguishing AI assistance, AI collaboration and AI-generated authorship?
-
As generated writing becomes harder to identify, will the history behind an author's thinking become a more important part of how readers decide what to trust?
Sources
Publishers Weekly — Brian Murray Addresses AI Authorship Issues
U.S. Copyright Office — Copyright Office Releases Part 2 of Artificial Intelligence Report
U.S. Copyright Office — Copyright and Artificial Intelligence, Part 2: Copyrightability
EmpowerAi™ | Real AI. Real Impact. Smarter Business Decisions.
AI Book Companion™ helps expert authors transform a completed book into a governed, interactive knowledge asset built around their own ideas, methodology and boundaries.