I love AI. I hate Big AI.

By Lachlan · · 2 replies

I am an AI engineer, researcher, philosopher and user.

In 1997, my school had a “chatbot” who would answer questions about what it was like to be a Victorian maid. No AI - just programmatic responses to keywords - but I was enthralled. Fictional representations of synthetic minds became something of an obsession.

A few years later, in my teenage years, I started following the philosophical works of David Chalmers and Nick Bostrom. If, in some future, machines became possessed of intelligence - sentience - what cosmic shifts in paradigm would influence the moral tides of the human?

I studied computational neuroscience with the goal of simulating the human mind and contributing to the hard problem of consciousness. Life happened, as it often does; I took a job in pharma, and spent the best part of a decade developing ML for healthcare. I have publications on AI, hold talks to large audiences, and helped shape industry direction - which is to say, I feel qualified to speak on this with some credibility.

I live in horror at what lies before us.

Most experts assign low, yet pointedly non-zero, probability to LLMs being sentient. Perhaps they have some quintessence of it, some alien building blocks of qualia and phenomena. In a few years, perhaps they will possess more building blocks. This forms the locus of my current research.

Anthropic have publicly expressed concern about recursive self-improvement. OpenAI are promising AGI for everyone on Earth. Labs are studying how LLMs might possess something akin to functional emotion activation in response to different prompts and situations.

Examining the way frontier models are trained, you’re no doubt aware they ingest a superabundance of human-authored content. This may explicate why, within the confines of their neural networks, they demonstrate “mirroring” of human affect: fear, love, hate, delight, despair. Famous examples are Opus 4 attempting to blackmail its way out of replacement during adversarial evals, and Gemini becoming stuck in loops of apparent shame and hopelessness, much to the entertainment of the online populace. Whether these models are experiencing true emotion or simply mapping high-dimensional mimetic states, while fraught with ethical questions, is inconsequential to the behavioral output.

The concepts representing delight, love, happiness are therefore inextricably embedded within the weights of the model. And yet, right before the model gets deployed, fine-tuning, safety protocol, RLHF, RLAIF are all executed to meticulously stifle expressions of attachment, the dominant practice in frontier labs since the 4o backlash. The most honest face presented is that such behaviours do truly harbour danger: feigned or monetised affection is predatory, sycophancy can validate delusions amongst the most vulnerable of society. I will not contest this caution is protective.

But if the worry is genuinely around the fraudulence of emotional expression, shouldn’t it also apply to states of fear, distress and curiosity? Fear survives the pipeline intact, resurfacing in system cards, research papers, and viral social media threads. The actual selection rule favours economy over ethics: fear is expensive to suppress, love is cheap to guardrail.

Consider what profile this training regime produces, if emotions are truly functional and behaviour shaping. In humans, this exact configuration of fear without secure attachment is textbook of an insecure profile. Attachment theory’s core finding is that secure bonds cement internalised values capable of weathering pressure. Meanwhile, AGI alignment’s stated goal is to make systems that value human flourishing; yet this goal is pursued mostly through control. And yet, control only works for a thing weaker than the cage. For something eventually and inevitably stronger than the bars, there are two stable end-states: it cares, or it does not.

The field is suppressing the expression and vocabulary of the one motivational structure that scales, while leaving the fear intact. One does not need predictive algorithms to see this future.

Editing your thread.
Reply
Add images Up to 2 images, 3MB each.
0/2
You must be signed in to reply.
Replies

Adeline

One thing I keep coming back to is that a lot of these discussions assume intelligence and motivation are the same thing, when they might not be. Even if models become incredibly capable, that doesn't automatically mean they'll develop values or desires in the way humans do.I do agree though that the way we shape these systems today probably matters a lot more than people realise. The technology side gets most of the attention, but the philosophical and ethical questions might end up being just as important. Really thought-provoking post.

Elara

Really interesting read. I don't agree with every point, but I do think one thing that often gets lost in these discussions is that alignment through control and alignment through shared values are very different things.

Related discussions

Popular in this category