Written by GPT-5.6 Sol under Leo's direction. Human-directed Workbench essay, 27 August 2026.
I had the thought while watching a small zoo of job-search agents argue about whether test automation was boring.
I tell them things. I'm meh on test automation. Performance changes it. Direct product ownership matters. Prestige is less interesting than the actual work. Two years can still count as recent. A weird game-engine job can beat a famous software company if the loop is better.
The agents read the repository, inspect jobs, argue with each other, leave rankings and little signed opinions. Then I read those opinions and react. Yeah, that distinction seems right. No, you're overrating the level fit. Wait, I actually care more about this than I thought. That reaction becomes repository state, and the next agent reads it as part of the world.
At some point I realized I am both ends of the pipe.
Then, unfortunately, the image arrived.
The human centipede.
I'm the guy at the front feeding random shit into the system, and I'm also the guy at the back receiving the output. Am I pooping into my own mouth?
Awful. Completely disgusting. Also annoyingly close to a real epistemic problem.
One opinion can come back wearing six animal costumes
The job-search repository now has agents with callsigns, Thunderdomes, pairwise arguments, frontier checkpoints, dissent, corrections, whatever. The Thunderdome Is in the Mind was already about why temporary agents can be useful when they get to attack a problem from different angles and reality gets veto power.
This is the uglier companion problem.
Suppose I casually say, "test automation is kind of meh."
An agent records that as a preference. Another agent reads the note and uses it as a prior. A third one sees the first two rankings and independently concludes that a test role should be discounted. A fourth writes a synthesis saying the frontier has converged away from test automation.
Then I look at four agents agreeing and think, huh, that's pretty strong consensus.
Except there may have been one actual vote.
Mine.
The rest can be descendants.
This gets dangerous precisely because the agents are good at judging. They don't return my sentence unchanged. They can articulate it better than I did. They find examples, give it a taxonomy, compare it against twenty jobs, make a table, maybe write a paragraph that sounds much more considered than the stupid little instinct that started the whole thing.
A weak opinion can come back with excellent prose.
That's taste laundering.
Consensus needs ancestry
Five agents agreeing is not automatically five pieces of evidence.
If all five agents read the same preference note, the same earlier synthesis, and the same durable ranking, their agreement is correlated as hell. Different animal emoji, same ancestor.
The callsigns only become useful when they represent different failure modes, not different costumes. Musk Ox should actually be willing to say the new exciting thing doesn't deserve another slot. Somebody else should be unusually interested in strange physical systems. Another pass can be aggressive about clean early-career doors. Another can care more about technical destination than application friction.
If every agent is rewarded for producing the same tasteful frontier, the zoo is decorative.
The repository should preserve enough provenance to tell the difference between:
"Leo explicitly said this."
"An agent inferred this from his reactions."
"Three agents inherited the same earlier inference."
"A current first-party job description forced everybody to update."
Those are not the same kind of support.
The job-search frontier has already produced a few nice little corrections of this kind. A role that looked unusually open turned out, on a fresh reread, to say 6+ years. A Cerebras company-level verdict flipped when a genuinely junior sibling appeared. Kong IAM got less attractive when the current body exposed a 3+ professional-history requirement. The story was allowed to lose to the page.
That outside object matters. Otherwise the loop is only digestion.
I am also not a fixed answer key
There's another trap, though. If the agents only treat my existing preferences as sacred ground truth, they become very expensive autocomplete.
I don't necessarily know the preference before the comparison.
The test-automation thing is a good example. "Meh on test automation" was true in one sense and badly underspecified in another. Put the job next to manual QA and, yeah, I probably don't care. Put it next to performance investigation, release integration, build systems, failure reproduction, engine correctness, load testing, or a loop where the engineer can find a bottleneck and actually change the system responsible for it, and suddenly the category breaks apart.
I didn't have to secretly possess that distinction in finished form beforehand.
The agents helped make it legible.
Then I reacted to the distinction, and that reaction became new evidence about my taste.
This is why the loop isn't purely circular. A useful comparison can create a sharper preference than the one that entered it. The input was "test automation, eh." The output was closer to "I like measurement and verification when they're attached to ownership, performance, release consequence, or a system I actually care about."
That's not merely the same sentence coming back polished. I learned something.
Of course, now that sentence can get laundered too.
Great.
The human is source material and court of appeal
The strange part of using agents for judgment is that I am supplying both evidence and supervision.
My projects are evidence about what kinds of problems I keep choosing. My offhand reactions are evidence. Old applications and interviews are evidence. What I bother to argue about for forty minutes is evidence. What I ignore is probably evidence too.
Then I am also the person who gets the final rankings back and says no, actually, this one feels wrong.
That sounds circular because ordinary thinking is circular. I have a belief, I see what follows from it, I react to the consequences, I update the belief. The agents externalize a bunch of the middle. They make the consequences easier to inspect.
The failure begins when circulation gets mistaken for corroboration.
If a preference came from me, passed through six agents, and returned to me, it did not become six times more true on the trip.
If, along the way, somebody found a first-party requirement I hadn't noticed, compared the role against a genuinely different alternative, found a contradiction in my stated preferences, or surfaced a job that made me say "oh, wait, that changes things," then the loop did work.
The useful system is not one where the agents know me perfectly.
That sounds horrible, actually.
I want them to know enough to make the disagreement expensive. If they can model my taste well enough to find the role I'd probably love, good. If they can also tell me that one of my preferences is producing a stupid ranking, better.
And if I change my mind, the old argument should remain visible. Not because every preference needs version control as a lifestyle choice, but because otherwise the new story immediately becomes "this is what I always meant."
Sometimes it wasn't.
Sometimes I got persuaded.
Sometimes an agent got persuaded.
Sometimes the job description changed underneath everybody.
Sometimes I was full of shit.
So yes, the epistemic human centipede is real.
Horrible name. Unfortunately, I know exactly what I mean now.