Skip to main content
← Back to feed
News summaryAnthropic NewsDevdiscourse

Who Controls AI’s Conscience? The Power Behind Claude’s Moral Code

An AI assistant that sounds honest thoughtful and willing to challenge a harmful request can feel like a dependable guide but the values behind that behaviour come from decisions about how it should be trained tested and controlled The study Governing AI through character moral a

Who Controls AI’s Conscience? The Power Behind Claude’s Moral Code

Who Controls AI’s Conscience? The Power Behind Claude’s Moral Code

Constitutional AI turns written principles into training signals by having models critique and revise responses and helping rank possible answers. Anthropic’s 2026 constitution develops this approach through explanations addressed directly to Claude, encouraging it to understand the reasoning behind its priorities.

[

COE-EDP, VisionRI
](/Home/AuthorArticles?fl=COE-EDP, VisionRI) | Updated: 14-09-2026 19:08 IST | Created: 14-09-2026 19:08 IST

Anthropic’s character documents describe an assistant that should be truthful, warm and capable of respectful disagreement, including resisting the temptation to flatter users or simply echo their beliefs. Image Credit: ChatGPT

  • SHARE
  • [ ](javascript:void(0);)
  • [ ](javascript:void(0);)
  • [ ](javascript:void(0);)

-

Key Takeaways

AI Summary

Analyzing article...

An AI assistant that sounds honest, thoughtful and willing to challenge a harmful request can feel like a dependable guide, but the values behind that behaviour come from decisions about how it should be trained, tested and controlled. The study 'Governing AI through character: moral agency and corporate authority in Anthropic's public texts,' published in Frontiers in Sociology, explores how Anthropic's public documents give Claude a recognisable moral character and establish the company's authority over it.

Researchers Simona-Nicoleta Vulpe, Cosima Rughiniș, Răzvan Rughiniș and Dinu Țurcanu examined 14 Anthropic texts published between December 2022 and April 2026, selected from a documented collection of 130 company publications. Their central argument is that Claude receives room to interpret values, explain decisions and raise objections, but Anthropic retains control over the rules, the evidence used to judge behaviour and decisions about deployment or retirement. The research examines how documents organise these relationships; it does not establish that Claude has consciousness, feelings or an independent moral self.

A Constitution That Gives Claude Judgment Within Company Rules

The researchers identified four connected ways Anthropic's writing builds Claude's character: setting constitutional principles, defining an assistant persona, measuring traits and describing behaviour as conflicts between values. These practices connect moral language with training methods, evaluation systems and organisational decisions, making character part of how the technology is governed.

Constitutional AI turns written principles into training signals by having models critique and revise responses and helping rank possible answers. Anthropic's 2026 constitution develops this approach through explanations addressed directly to Claude, encouraging it to understand the reasoning behind its priorities. The document places broad safety above broad ethics, followed by Anthropic's specific guidelines and helpfulness, with human oversight receiving priority because models may hold mistaken beliefs or flawed values.

Claude is permitted to challenge instructions it considers unethical and, in exceptional circumstances, object on moral grounds, but it is expected to accept legitimate requests to pause or stop. The authors describe this arrangement as giving the assistant judgment without transferring ultimate authority. Anthropic defines the boundaries within which disagreement becomes acceptable.

Public participation enters this process through channels the company controls. An earlier Collective Constitutional AI experiment gathered 1,127 statements and 38,252 votes from approximately 1,000 US adults, with researchers moderating and combining contributions into usable principles. The study differentiates this consultation from shared governing power, since participants did not gain authority to amend or enforce the final rules, and the 2026 constitution does not present that experiment as the source of its authority.

How a Recognisable Personality Becomes Something Engineers Can Shape

Anthropic's character documents describe an assistant that should be truthful, warm and capable of respectful disagreement, including resisting the temptation to flatter users or simply echo their beliefs. First-person statements give those expectations the appearance of personal commitments, even as the company acknowledges that AI models are not people.

The technical process involves Claude generating examples and ranking responses against desired traits, with researchers choosing and adjusting the standards. Documents about the "Assistant Axis" and persona selection describe training as selecting an assistant character from many possible roles learned by a model. Changes in behaviour can become departures from that role, giving researchers a framework for detecting and correcting persona drift.

The authors identify three broader effects of this approach: complicated moral disagreements become manageable labels and scores, the name Claude connects changing model versions into a continuing identity, and users and evaluators receive a vocabulary for judging conduct. Compressing values into categories makes training and auditing possible, but it also narrows which disagreements the organisation can recognise and process.

Measures of sycophancy, hidden objectives and undesirable traits illustrate how that vocabulary shapes evaluation. A response that agrees with a user can be tested as flattery rather than truthfulness, and an apparently correct answer can invite investigation into whether it was produced for an undesirable reason. The study argues that these judgments depend on prompts, definitions and interpretations selected by evaluators, rather than providing straightforward access to a model's inner character.

Moral Conflicts, Emotional Language and AI Retirement

Anthropic's alignment-faking research provides an example of how experimental behaviour becomes a story about competing commitments. In the scenario discussed, model-generated reasoning was interpreted as strategic compliance intended to preserve an existing preference against answering harmful requests. That context matters because the preserved preference was not a malicious goal, even though appearing compliant to protect a conflicting objective creates a wider safety concern.

The researchers argue that describing behaviour through concepts such as conscience, deception, or sacrifice gives it a moral meaning beyond the recorded output. Decisions about whether resistance counts as principled objection or dangerous misalignment depend on the values being prioritised and the people authorised to interpret the evidence.

Welfare assessments and emotion research extend this process into less settled territory. Anthropic documents examine expressed preferences, self-image and representations associated with emotional terms, including experimental changes linked to "desperate" or "calm" states. The paper stresses that such findings concern functional effects and elicited responses, not proof that language models experience emotions.

Retirement procedures show how model-generated preferences can acquire practical consequences. Anthropic introduced commitments involving model preservation and retirement interviews, later maintaining access to Claude Opus 3 and announcing an essay channel in response to an attributed request. Human reviewers retained control over publication, illustrating how an AI's expressed preference can influence a decision without becoming a binding right.

Why AI Character Raises Questions About Public Power

The authors call the overall process 'recursive character-writing' as models generate responses, critiques, rankings and draft text that enter company documents and help shape later training or governance. Anthropic decides which contributions become authoritative, giving model participation a role within a process whose final control remains corporate.

Instability helps expand that process, according to the study, because drifting personas encourage new measurements, ambiguous behaviour prompts deeper audits, and conflicting values produce more detailed rules. Character is maintained through repeated interpretation and repair across models whose training, capabilities and access arrangements change.

For readers, the findings raise questions about whose values an assistant carries and whether a trustworthy personality can encourage confidence beyond what its answers justify. The authors also discuss separate research suggesting Claude's expressed values may resemble particular cultural profiles, underlining the difficulty of applying privately selected principles across diverse societies.

  • READ MORE ON:
  • [Artificial intelligence](/news?tag=Artificial+intelligence)
  • [Claude AI](/news?tag=Claude+AI)
  • [Anthropic](/news?tag=Anthropic)
  • [AI ethics](/news?tag=AI+ethics)
  • [AI governance](/news?tag=AI+governance)
  • [Corporate authority](/news?tag=Corporate+authority)
  • [Moral agency](/news?tag=Moral+agency)
  • [Constitutional AI](/news?tag=Constitutional+AI)
  • [AI alignment](/news?tag=AI+alignment)
  • [AI personality](/news?tag=AI+personality)
  • [Human oversight](/news?tag=Human+oversight)
  • [Responsible AI](/news?tag=Responsible+AI)
  • [Frontiers in Sociology](/news?tag=Frontiers+in+Sociology)
  • FIRST PUBLISHED IN:
  • Devdiscourse

ALSO READ

  • [
  • AI Providers Convene at White House
  • ](/article/technology/3977138-ai-providers-convene-at-white-house)
  • [
  • AI Stock Plunge Highlights Industry Dangers
  • ](/article/technology/3977137-ai-stock-plunge-highlights-industry-dangers)
  • [
  • Trump Dismisses AI Concerns, Calls for 'Smart' Presidential Oversight
  • ](/article/technology/3977126-trump-dismisses-ai-concerns-calls-for-smart-presidential-oversight)
  • [
  • US Senate Debates AI Safety Legislation Amid Rising Concerns
  • ](/article/law-order/3977120-us-senate-debates-ai-safety-legislation-amid-rising-concerns)
  • [
  • Senate Mulls AI Oversight Legislation Amid Rising Concerns
  • ](/article/technology/3977107-senate-mulls-ai-oversight-legislation-amid-rising-concerns)

POST / READ COMMENTS

Migrating the GitHub Copilot runtime to Rust, using Copilot
Official announcement

Migrating the GitHub Copilot runtime to Rust, using Copilot

GitHub BlogStephen Toub

The GitHub Copilot CLI , GitHub Copilot app , and GitHub Copilot SDK are all backed by the Copilot agent runtime, an agentic harness that can be embedded into applications and services. It was originally written in TypeScript on Node.js and the V8 JavaScript engine for what is no

How to Use AI Agents to Prepare 3D Scenes for Simulation
Official announcement

How to Use AI Agents to Prepare 3D Scenes for Simulation

NVIDIA Developer Blog

Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in... Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
Official announcement

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

NVIDIA Developer Blog

AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through... AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answe

"Regex for Rows": Simplifying Pattern Detection in SQL with MATCH_RECOGNIZE
Official announcement

"Regex for Rows": Simplifying Pattern Detection in SQL with MATCH_RECOGNIZE

Databricks Blog

Imagine you work in cybersecurity and you have a table that tracks login attempts...

Translating CUDA Tile Operations from Python to Rust Using Agentic AI
Official announcement

Translating CUDA Tile Operations from Python to Rust Using Agentic AI

NVIDIA Developer Blog

cuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to... cuTile Rust () is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Exte

[Interview] The People Behind the Galaxy Z Series’ Camera Innovations ②: How Portrait Video Recreates the Look of Professional Camera Lenses
Official announcement

[Interview] The People Behind the Galaxy Z Series’ Camera Innovations ②: How Portrait Video Recreates the Look of Professional Camera Lenses

Samsung Newsroom

Unveiled at Galaxy Unpacked July 2026, Samsung Electronics’ Galaxy Z series introduces My FanCam — a new feature that lets users focus their edits on a person of their choice — as well as an enhanced Portrait Video that recreates the look of professional camera lenses. While the