Tag Archives: artificial-intelligence

Academic publishers have set some guidelines about AI use and, well, it’s a start

I like writing about the scourge of generative AI in the arts – obviously – and it seems like there’s new news almost every day.

I’ll make this quick:

In the new Publishers Lunch is a report on how some publishers have developed guidelines and methods of screening papers and books for AI use. (You’ll recall that the suspicion of AI has roiled the mainstream publishing industry lately, causing book deals to be canceled and, possibly, author careers to be stymied.)

Academic publishers cited in today’s Publishers Lunch have tried to come up with a constructive method of screening works submitted to them for use of AI. That’s because like a lot of publishing, academic publishers are being swamped with submissions of AI or AI-assisted slop.

The guidelines, which build on the use of AI-detection tools, set thresholds for when AI use is detected.

Some of the publishers have established four guidelines:

  • Assisted – AI helped you brainstorm or research but the writing is yours. Meaning, if you asked AI to find the relevant information for you and then you read those source materials and wrote from them, that’s a mechanical assist. It’s just AI working as a search engine.
  • Collaborated – AI drafted parts while you rewrote, verified, and made it your own. Meaning, if the facts, ideas, etc. came out of the AI model itself (e.g. you asked it explain something to you) and wrote from its answer instead of the source materials, that is AI collaboration and should be disclosed.
  • Generated & reviewed – AI produced the artifact, but you read every line, checked the facts, and stand behind it. To be clear, you hold responsibility for its integrity and accuracy.
  • Generated & unreviewed – AI made it and you didn’t really check. This is not acceptable AI use for PUP staff ever: internal, external or under any circumstance.

Publishers Lunch says that for at least one publisher, if a submitted work meets the last two guidelines, it’s rejected. That’s because that particular publisher correctly argues that the use of that much AI couldn’t be copyright-able because, of course, AI is based on the work of others.

I’m skeptical, but at least somebody is setting some guidelines somewhere.

(Image at the top of this post is from the wonderful, funny AppleTV Plus series MURDERBOT, based on the series of novels by THE MURDERBOT DIARIES by Martha Wells.)

Report: Amazon is buying old books to scan for AI and destroy. Yes, this surprises no one.

I had intended to follow-up my weekend post about getting my books in a local bookstore by suggesting what authors might do to achieve the same thing, but I’ve delayed that for the moment. First things first, and just like a house on fire, the dual plagues of AI and Amazon get our attention at the moment.

For a few weeks now, book lovers and people who hate AI have been despairing at reports about how companies that steal, buy or pirate books to train generative AI – thus putting creative people at risk of losing their livelihoods – have been buying and scanning books. The revelation comes out of the landmark class action lawsuit Andrea Bartz and other writers filed against Anthropic. Bartz and company won a $1.5 billion settlement on behalf of many, many writers whose work was stolen by the AI company.

Details of Anthropic’s practices came out in the discovery process, and they’re horrifying: The company was buying hundreds of thousands of copies of books, some of them older and out of print, and they’re cutting off the spines and feeding the pages into AI-making programs. The books are utterly destroyed. The judge in the Bartz case ruled that this practice was legal because the company had bought copies of the books, not stolen them. (They’re using older books to ensure no AI writing gets fed into the AI consumption process, of course.) Videos of this process – chopping spines off books and the ultimate shredding of the pages – are disgusting and disheartening.

And of course Anthropic isn’t the only AI company doing this.

Today, thanks to an ingenious investigation by 404 Media, we also know that Amazon is doing it. Maybe not a surprise – Amazon and its owner are widely regarded as awful; most of us who do business through them do so only reluctantly, or at least that’s how I feel – but the confirmation from 404 Media, which was also shared by Publishers Lunch this morning, is damning,

This is the lead of the 404 Media story published today:

Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.

(I’ll link to the 404 Media story at the end of this piece.)

404 Media cleverly hid an AirTag-type tracker in an old book that was to be sold to Amazon’s book-buying operation. 404 was able to track the book as it and many others were shipped around the country and ended up in Las Vegas, where it went to an Amazon warehouse used by an operation called VGT3. That’s the operation’s logo above. Charming, huh?

Amazon didn’t really address the process or its value to anyone but Amazon in its response to 404.

(404 had, in July, reported on an increase in sales of old books and speculated that the increase in sales was due to AI-generating companies buying books it could scan and destroy. And again, I’ll note that the books being chopped and scanned and shredded were published before 2022 or so to AVOID THE WRITING BEING TAINTED BY AI.

As 404 reported: These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn’t say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer.

In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.

404 Media granted the bookseller anonymity because they worried sharing this information would harm their business.

404 found online comments from Amazon employees about the VGT3 facility that virtually all they do there is scan books.

Many thousands of books are being bought, scanned and destroyed by Amazon at that facility and probably others. And that doesn’t even account for other companies that are doing the same.

And as 404 reported:

That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.

I already didn’t feel good about Amazon, where a large amount of my work can be purchased. Their practices toward employees and authors and the sheer mass accumulation of wealth by its owner were long-known and queasy-making.

Here’s another reason, thanks to the reporting by 404 Media and dissemination to the publishing industry by Publishers Lunch, to dislike Amazon.

Here’s the 404 Media story:

https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility

With latest AI accusations, book publishing industry continues to melt down

The book publishing industry continues to work hard to self-destruct.

Everybody who writes about book publishing and particularly the harm wrought on books, writers and largely unsympathetic publishers is reeling today and writing about this, so while I’m not reeling I am writing – and slapping my forehead.

And yes, I’m being mean with that headline, because it’s not solely, or even primarily, the publishing industry trying to self-destruct. The big tech AI firms, which make generative AI possible, steal everything that isn’t nailed down, whether it’s writing or art or even something mundane that can be used to generate email.

And yes, I feel strongly about this because, last time I checked, two of my early true crime books were stolen to train AI book theft programs. “Wicked Muncie,” my co-written 2016 true crime book, and “Muncie Murder and Mayhem,” from 2018, turn up in a search of my name on the database of pirated books as provided by The Atlantic.

Without going into too much detail, in the course of this week, at least two “big” books have been withdrawn by their authors, stillborn by their publishers or otherwise fingered for having artificial intelligence elements. This can include AI writing or AI editing. A lot of people are giving the authors the benefit of the doubt and I can understand and appreciate that. Others are crediting the publishing companies for throwing a flag on the books because it’s not clear, despite their “we tested this book and our tests show it was written and/or edited with AI” claims.

Everyone from Publisher’s Marketplace to newsletters by Kathleen Schmidt (Publishing Confidential, an essential newsletter) and others have recounted the allegations and resulting controversy.

As Schmidt says, “Things can’t go on like this.”

Why do we care about this? Well, besides the fact that AI is an absolutely unnecessary tool in the creative arts, and besides the fact that data centers used to create AI products and portals are horribly destructive to the environment because they use rampant amounts of water and often run on heavy-polluting diesel generators, the fact is that the use of AI in creating art of any kind is destructive to creators, authors, artists, editors and, as much as we don’t sympathize with many of them, publishers.

As Schmidt notes, “a $2.5 million, two-book deal that was acquired by Minotaur (Macmillan) in a 14-way auction, (was) pulled by the author’s agents due to concerns about AI use in writing .”

This followed on the heels of another book and author questioned. As Schmidt reported, “Earlier this week, the author H.M. Smith was accused of using AI to write her new book, Daggermouth. In that case, Pangram returned the book as 60% AI. It is currently ranked at 73 on Amazon. Before that, Hachette canceled the book Shy Girl because Pangram returned it as 78% AI. What’s notable is that in each case, it was an individual from a media outlet feeding book files to Pangram without the authors’ permission. Welcome to the new world of “Gotcha!”

Schmidt is among those who note that programs that use AI to supposedly detect AI are highly probable to find AI where it doesn’t exist. There have been reports of long-ago-published books that, when tested by AI, come back with a false positive for the use of AI. Schmidt notes that she ran some of her newsletters through Pangram and the result was that they were 80 percent AI. And, of course, she’d written every word.

Here’s where I get to the self-destruction part. As Schmidt notes: Whether you admit it or not, there is A LOT of jealousy in the publishing world, so it wouldn’t surprise me if someone who disliked author X decided to use Pangram on their work and made the results public. This is not the publishing world I want to live in.

As Publishers Lunch/Publishers Marketplace notes, Sources tell PL that more deals are being cancelled behind the scenes following the discovery or suspicion of AI — one publisher noted that multiples imprint have had to cancel a book over this — including books that originate in the self-publishing sphere and are then picked up by Big Five publishers. This usually happens quietly and under the public radar, and is only brought to light by fans online.

There’s plenty of blame to go around here. There are apparently a ton of AI-written or AI-assisted books on Amazon. People who want to be authors but are unable to produce a book themselves or unable to hire qualified human editors to read their work – hell, just find writer friends who will edit your manuscripts and maybe pay them or repay them in kind – are causing readers to buy and read dreck and causing publishers to suspect everybody.

Before this latest AI controversy, I was going to write about another AI controversy: the abysmal practice, which became known in the wake of author Andrea Bartz’ $1.5-billion-winning class-action lawsuit against the AI company Anthropic, of Anthropic and other AI pirate companies, of buying copies of books and cutting off their spines in the process of feeding them into AI readers then shredding what’s left. (By the way, you should follow Andrea Bartz on social media. She’s become the face of authors fighting back against AI, but she’s also a great writer of thrillers.) This slice-and-dice process was allowed by a judge because once you’re bought a physical copy of a book, you can do as you want with it. Including destroying it.

So yes, Schmidt is correct. This mess can’t continue like this. It will, of course, as long as AI companies, publishing companies and some authors continue to make money. But in the meantime, the environment is suffering, authors and editors are suffering and writers whose work was stolen or must compete with AI slop are suffering.

AI, Michael Caine and Mr. Potato Head, together again

I mentioned on social media in the past couple of days that I got the biggest reaction I’ve ever received on LinkedIn after posting that i’d unfollowed a LinkedIn connection after seeing them post touting AI.

Now this isn’t unusual on LinkedIn, where a lot of people have some financial investment and a vested interest in seeing AI succeed.

I posted this:

Unfollowed and cut my connection with an AI user.

AI is a destructive force. It kills jobs, creativity and the environment, all for profits for billionaires who don’t care about any of us. And when the AI bubble bursts, the economy will suffer.

If you’re an AI user, go ahead and unfollow me and cut our connection. You might as well, because when I see you touting AI, I’ll do it.

This brought responses from people I actually know and some that I don’t, taking the approach of warning me that: I’ll never get ahead without AI – similar to Reese Witherspoon’s “don’t get left behind” warning to all her girlies out there – and that I’m already using AI but probably don’t know it and that I’ll be using it in the future because everybody will be.

My answer was more polite than “bullshit” but that was the gist of it.

Anyway, the whole thing was amusing and good for engagement and I’ll probably go to that well again, despite the dire warnings – all, without a doubt, from people who have something invested in AI or at least hope to make a buck in it – popularizing the notion that all of us who have been writing email, writing books, etc., for decades WITHOUT the assistance of AI have apparently lost the ability to do so unless we rely on the processes that are killing the environment and killing jobs just so some poor schmuck can imagine themself as an author or have a “girlfriend who won’t say no.” (See my recent post on the topic of the AI girlfriend.

Then this morning, Publisher’s Marketplace reported on an “AI voice company” that plans to release an audio version of “The Odyssey,” timed to coincide with the Christopher Nolan movie, that will be narrated by the AI version of Michael Caine’s voice.

I can guarantee you I will live the rest of my days without listenng to that.

Those of you who know I have an absurd sense of humor know that the final paragraph of the article was my favorite:

ElevenLabs primary business is creating synthetic AI voices and text-to-speech audiobooks. They have partnered with Spotify to produce audio for self-published authors, and digital distributor Bookwire. Their Iconic Marketplace allows brands to license famous voices for AI-created content, in partnership with the celebrity or estate. Currently available voices include Dr. Maya Angelou, Judy Garland, David Hasselhoff, Laurence Olivier, and Mr. Potato Head.

So I want to know, did the company license the AI rights to the voice of Don Rickles, who voiced Mr. Potato Head in the “Toy Story” movies? And if so, why not just say their available voices included “Don Rickles as Mr. Potato Head?”

So your books have been pirated to train AI …

As a writer, I find AI an intensely bad thing. Yeah, it’s momentarily distracting and amusing to be scrolling through social media and see what are obviously AI-generated images of a horrible, horrible person licking the feet of an equally horrible person, or even to see some fanboy’s imagining of what a Justice League movie would look like if it were made in the 1960s.

Then you realize that this is AI and valuable natural resources are being used to run servers that create these images. Not to mention that real, actual artists – and in the case of the written word, writers – could be put out of work by this.

I first had some foreboding realizations about the effect AI might have on my work a few months ago when I went looking for one of my pieces for CrimeReads, so I could post a link to it, and realized that Google AI had generated bullet points of my articles. Why would someone need to click through to CrimeReads when they could just read the AI interpretation of what the site’s writers had written?

I was aware that some writers were saying they believed entire books of theirs had been used for AI training.

I was concerned about that because I know a lot of writers. I thought no one would possibly pirate and upload my little true crime books. Who would need that?

Then, on March 20, the Atlantic published a story about the 2-million-plus books, articles and scientific papers that have been added to Library Genesis, or LibGen, which is what Wikipedia calls a “shadow library” of file-shared work, including work that is not available digitally.

One of my writer friends said a couple of her books were there. Another had 15 of her 19 books pirated on LibGen.

The Atlantic offered a real public service that allowed readers to search to see if their work, or the work of someone they know, was uploaded to LibGen for AI training.

Here’s a screen shot of my search results using the link in the Atlantic.

I found two of the four true crime books I co-wrote with Douglas Walker on there, using the Atlantic’s search engine.

I later found what purports to be LibGen’s own search portal and could not find these two books. Had they been taken down in the meantime? Was there some mistake? It seems hard to imagine that the Atlantic got that wrong. Based on that article, I saw dozens of writers, some of whom I know, posted that they had also found their books on the site.

It’s unclear what to think about what’s there and what’s not, but there’s no question that pirated work hurts writers and publishers who might not be able to sell copies of books if people can get them for free. Since shit flows downhill, that trickles down to harm for writers, that’s for sure.

We’re all still figuring this out. It’s been pretty clear for a while that AI-generated art and writing is bad for the planet – servers use a lot of water to cool to create AI – and bad for writers. I suspect it’s also bad for consumers, but then I was never one to snap up pirated books and art and have been pretty skeptical of that inclination.

Brave new world, hell. This seems like a very cowardly ploy.