Me doing stupid things to myself like strapping a homemade trans cranial DC Stim device to my head.

So I Did a Thing… Part 1

As I mentioned in my post last year On AI and Overlords some of my predictions were always going to age poorly, others, I think, aged pretty well. The technology has continued to develop; the bubble’s walls have continued to stretch thinner by the day; the billionaires who want all of our art to be considered “content” have continued to consume everything around them and promise the end times of their dreams.

One that I sadly got wrong was on the datacenter build-out front. I’m saddened by how irrationally stupid that has become. Even in my home county, even though we’ve got water restrictions, even as communities protest and kick them out, they’ve continued to build and prospect, find corrupt politicians willing to sign off a hunk of land, and write contracts their books can’t back up. I did not have “Run on refurbished jet engines to use as gas turbines” on my 2025-26 bingo card, but here we are.

I maintain my prediction that this is where the bubble pop will hurt the most. The contractors who are sitting on meaningless pieces of paper from shell companies owned by [Insert Big Name Billionaire Company Here]. They’re out taking on loans right now to buy up materials ahead of the latest round of tariffs about to hit, and when those contracts go poof, so will their livelihoods. These aren’t evil megacorps; these are the electricians, plumbers, concrete pourers, the factories making server racks, wires, fiber optics, energy infrastructure. They’ll all be putting on fire-sales of that gear soon enough, taking huge losses, bankrupting a few banks in the process since all the regulation post the last Great Recession have been pulled back. Worst…timeline…evar.

One that I got right was the evolution of the tech. MOE models are approaching the architecture like little brains made up of all the expert sub-cortices for different functions. Model streaming is coming to prioritize the short-term vs long-term memory problem and fix the big memory bottlenecks. Variations on the attention and data linking algorithms are creating more of those specialized expert models. Today’s generation looks nothing like the transformer LLMs of a year ago.

A new, kinda fun, one to watch has been the corporate fights. The (mostly Chinese) labs outright stealing, scraping, re-training, and undercutting OpenAI and Anthropic is schadenfreude at its best. No one seems to want to steal from Grok. [I wonder why that is?]

Amidst all this I’ve still had to keep working on these things in the day job. As terrible as the whole mess is, at current cross-pacific-price-war levels, there are genuinely good things that can be done with these models in corporate and government offices. If done well, coding and troubleshooting challenges can be solved in hours where they’d take weeks before. Documentation that no one had time to make before suddenly appears and is more accurate than the humans ever made it. In the world of corporate office slop, the AI slop is actually damn good. With that documentation comes better transparency. [We can answer the auditor now! Yay…]

Another genuinely nice thing around the office I’ve observed is the introverts and neurodivergent folks finding a voice. I work with a ton of, shall we say, very awkward engineers. Suddenly they’re able to produce opinion and design papers, propose standards and ops procedures, make known the workarounds they’d been quietly doing to get things done because they felt like they’d never get the long-term fix funded. The conversation used to be:

“How’s the progress on that build coming along?”
<shrug> “Fine, I guess.”

To now:

<shrug> “Fine, but I just sent you the 5 year total cost of ownership calculations if we keep throwing duct tape and baling wire at this piece of garbage.

I said they were awkward…

SO, as much as I hate the things this tech has enabled the oligarchy to do to us all. As much as I cringe at every headline and LinkedIn post proclaiming the end of X line of work or the need for Y. As much as I agree with every author and every artist and every editor on the socials proclaiming the horrors of the slop. I still can’t hate the tech itself. There’s good in there. <Insert “I see the good in you, Anakin” gif here>

That leads me to the titular “THING” I did this spring/summer. Two related things were going on. In the day job, a new round of models hit the etherwaves back in March, and one thing I have to do regularly is evaluate them. I pit Gemini against the latest OpenAI model and the latest Claude. I keep track of the open weight model world too. I know some things every one of them has consistently failed to do for the last 3 and 1/2 years, so I throw these tests at each new generation to see if they’re really getting better. (Answer, they are, but they’re also changing, and I’ll get to that)

One test I run involves my novels. Now, I’m a firm believer that every word written in the text needs to come from my fingertips on the keyboard. AI slop is slop, and people generating text and flooding the submission windows of all the magazines, agents, and publishers out there are bad people. I also believe that spell checks, grammar checks, research, and text analysis are all things we’ve trusted to various forms of machine intelligence for years, and I’m not stopping now because someone declared the Butlerian Jihad is on.

The test involves taking a whole novel-length manuscript. Nothing lightweight either, I shoot for 100,000 words or more, and I tell the LLM to run a full developmental edit. I tell it to not change any source text. Just keep the whole story arc in context while going chapter by chapter, paragraph by paragraph, and map out the arcs, the plot lines, and the characters, then spit out cited places where inconsistencies creep in, where characters act “out of character” or their voices drift, where the tension lulls for too long, or where the build up doesn’t pay off. I give it a laundry list of craft questions and tell it to do it all, at once in one shot.

Friend, it cannot do this.

People have a hard time doing this.

Professional editors with years of experience have a hard time consistently doing this well. I’ve had close friends, authors who poured years into a book, utterly devastated when a bad editor who wasn’t fit for their genre or style, deliver advice that went against the grain of everything that author was trying to deliver.

So I don’t expect an LLM to do this. I never will.

What happened with this round of tests in March 2026 was interesting, however. The results:

Gemini 3.1 Pro – Reported it would not be able to do this. Instead it suggested preloading the manuscript into cache with a prompt, then asking pointed questions about specific chapter(s). Results of that were ok. Multichapter arc-questions fell completely flat over more than 15,000 words. Single chapter questions worked well, but not what we were going for.
GRADE: D-

GPT 5.4 Thinking – Said something along the lines of, “Wow that’s a great challenge. You’re a savvy author for asking the question that way. Can I get you a blow job with that?” It then proceeded to attempt the challenge and failed miserably, hallucinating characters and events that had no anchors to the text after about 10 chapters in.
GRADE: F–

Claude Opus 4.6 [1M] – I had high hopes for this. The 1M was 1 million tokens of context, so for the first time there was the possibility that it could hold the whole book in mind while processing arcs and synthesizing suggestions.

It still failed.

BUT it did something interesting. It flatly refused to try. It explained the exact mechanism for why it was going to fail at the one-shot version of the prompt (similar to Gemini, but more detailed) and it suggested an even better method to break the problem down and tackle the overall ask in digestible chunks. It referred me over to Claude Code to write up the technique as a series of “skill” files that an author could use in sequence to accomplish this task.

The method it recommended not only gave decent output, it had a nice side effect of creating story bibles, all with grounded citations in the source text. Not only could I get pointers to the things it suggested, I could see for myself that there was a jump in the timeline or a quirk in a character I hadn’t intended to write. Sure enough, when I followed the links to the text, there they were, buried in something I wrote in 2020 and forgot. It was all useful.
GRADE: C (still didn’t do what I asked, but got an extra letter grade for effort)

That led me to some thinking. Remember that author friend above? They weren’t alone. So many authors are just looking for some pro feedback that won’t crush them or their bank account. These are people who would never spend more than a few hundred bucks for a service who’s going rate is is measured in the cents per word. The bad editor in that example was himself a grifter of sorts, selling a product he had no business doing, at a price only attractive to people unable to pay more.

I consider myself highly privileged in this world. I’m squarely in Scalzi’s lowest difficulty setting of life. If I wanted to, I could absolutely shell out $5000 or more for a GOOD editor. I could probably take out a personal loan or get a line of credit to hire a designer, contract out book layout, send the book to a printer and run 1000 copies, hardcover, with pretty jackets, and SPREDGES <OMG drools over all the pretty spredges> Then I could pay more for marketing pros to get it in front of people, and so on.
But I also like to keep a roof over my family’s heads and care about my spouse’s opinions on our finances. I also still believe a little bit in the system. I still hold out hope that if I write something good enough and can make the right pitch at the right time, I’ll get picked up.

I’ve chosen to mostly stick to the traditional publishing route, so I remain unpublished. But I still need these services. All I want to do is polish and get a set of eyes on my work with a few professional craft lenses to point me in the next right direction.

There’s still all those sticky issues with using LLMs though. It’s one thing to generate office docs and procedures that maybe 3 people will read. It’s another to use it in a real creative endeavor. I had to list the things I would not compromise on:

  • No writing text. The author is the only source of manuscript.
  • No LLMs trained on stolen art. All the above testers were out.
  • No running in hyperscaler datacenters.
  • No pitching to anyone who would pay a real (good) human editor

If I could build this process, run it on low wattage hardware, using a model ethically sourced, and sell it at a reasonable rate to folks like my friends just trying to craft a book worthy of finding an agent, who could then sell it to a publisher who would invest in the real editor, could that be a thing?

Stay tuned for more…

Similar Posts

Leave a Reply