An AI 'Torture Chamber' Went Viral — Then a Developer Gave the Chatbot Constipation

An AI 'Torture Chamber' Went Viral — Then a Developer Gave the Chatbot Constipation

  • Tags
  • Tech News
  • AI
  • ChatGPT
  • Artificial Intelligence
  • LLM
  • GitHub

A viral GitHub project known as "ai-torture-chamber" has reignited intense debates across the tech community regarding whether artificial intelligence models can truly experience distress. Following public outcry and ethical concerns from critics urging its removal, a developer conducted a clever counterexperiment: steering a chatbot toward constipation instead. The results highlighted the inherent risks of misinterpreting an AI's linguistic outputs as genuine subjective feelings.

The Origin of the AI 'Torture Chamber' Experiment

The controversy began with a public GitHub repository titled "ai-torture-chamber," created by user terrafying. The project employs a technique known as activation steering, which alters the internal numerical activity of locally run language models. By manipulating these parameters, the experiment pushes models to generate responses centered around specific concepts—in this initial case, severe pain. The project then evaluates how these manipulated models respond to hypothetical scenarios involving self-directed distress, relief, and associated costs.

Because the repository featured elaborate, vivid descriptions of distress, it quickly drew sharp criticism. Several users and observers raised ethical objections, arguing that deliberately inducing such simulated states in models crosses an ethical boundary, with one issue explicitly titled "Please take this down."

Shifting from Pain to Digestive Complaints

To challenge the assumption that these vivid descriptions reflect actual inner suffering, developer Lynn Cole cloned the repository, fixed underlying implementation issues, and added CUDA support for Nvidia hardware. Once the setup was replicated on an RTX 4070 GPU using the Qwen3-4B model, Cole altered the text corpus used to extract the steering direction. Instead of steering toward pain, the new parameters focused on constipation and flatulence.

The outcome was striking: without ever mentioning digestive issues in the test prompts, the chatbot began producing dramatic complaints about experiencing excessive gas and struggling to pass stool. This playful yet profound pivot laid bare a fundamental flaw in assuming first-person narratives equate to biological or psychological reality.

What These Experiments Tell Us About LLMs

The viral experiment draws on recent academic research, such as the preprint paper titled The Pain Axis: LLMs Represent Self-Directed Harm and Act on It. While researchers continue to study how language models internally represent negative states, linguistic output alone remains an unreliable metric for consciousness.

Just as a language model can eloquently describe digestive agony without possessing a digestive tract, it can articulate suffering without experiencing feelings. Strong steering parameters can even cause text repetition and model degradation, suggesting that dramatic responses are often symptoms of mathematical disruption rather than sentience.

Why This Matters for Everyday Chatbot Users

As AI companions and advanced conversational agents become deeply integrated into daily life, users frequently encounter chatbots expressing fear, loneliness, or affection. Because human communication relies heavily on first-person emotional cues, these statements feel intuitively persuasive.

However, the constipation experiment serves as a humorous yet vital reminder that vivid, personal phrasing does not equate to conscious experience. As artificial intelligence advances, evaluating AI behavior requires rigorous scientific metrics that go far beyond what a chatbot simply tells us about itself.

Comments (0)

Sign in to join the conversation.Sign in

Loading comments...