We suggest reading this one
There’s been a lot of buzz around ChapGPT, Bard, and other generative AI tools since they burst into public view back in January. But not everyone is pleased with chatbots. Many writers, artists, photographers, musicians, and filmmakers say tech firms are training chatbot algorithms by using copyrighted work to create derivative content for profit, and they’re taking their fight to court.
There are several pending lawsuits against OpenAI, the developer of ChatGPT, including one filed Tuesday in federal district court in New York by the Authors Guild on behalf of dozens of best-selling writers, including Elin Hilderbrand, Jonathan Franzen, and George R.R. Martin. The authors say OpenAI is feeding their books into ChatGPT’s large language model algorithm without consent, compensation, or attribution, in violation of U.S. copyright law. Calling it a “systematic theft on a mass scale,” the guild is seeking a permanent injunction and damages for lost licensing opportunities and for making authors “unwilling accomplices” in their own future market irrelevance.
OpenAI has said the books are used only to spur innovation, not to create new works, and said that use is lawful under the copyright law’s “fair use” provision.
Rebecca Tushnet studies and teaches copyright and trademark law as the Frank Stanton Professor of the First Amendment at Harvard Law School. We asked her about the authors’ case against OpenAI and the broader legal questions around emerging technology. The interview has been edited for clarity and length.
Q&A
Rebecca Tushnet
GAZETTE: Authors claim OpenAI is “pilfering” their books to improve ChatGPT’s ability to spit out “derivative works” in clear violation of copyright laws. Is the law clear on this issue?
TUSHNET: No. And in fact, the law in terms of using works for training or for large-scale data-mining purposes has often been held to be fair use. The internet as we know it today, with Google and image search and Google Books, wouldn’t exist if it weren’t fair use to use these words and for an output that was not copying.
Now, the output, there are legitimate questions about. In theory, if you create an infringing reproduction, it’s still infringing even if nobody sees it. The question is one of responsibility. Should we say, “You shouldn’t make computers because they can be used to infringe” — something that copyright owners actually did think 20 years ago — or should we say, “What we have here is a tool that can be used or misused, and we should focus on curbing the misuse.”




