← Back to writing

Technology & Society · October 11, 2026

The Battle Over AI Values Has Already Started

Anthropic's effort to train Claude's character shows that the battle over AI values is already a question of power, accountability, and public trust.

AI companies publish rules for model behavior. Anthropic takes a broader approach with Claude. It trains Claude not only to avoid harmful answers, but to display a designed set of behavioral dispositions that Anthropic calls "character."

Here, "character" means a relatively stable set of trained behavioral dispositions, such as honesty, caution, curiosity, and willingness to object. It does not mean a demonstrated inner self, consciousness, or autonomous moral agency.

Character guides behavior when rules conflict, are incomplete, or leave room for interpretation.

Anthropic turns values into training signals. Its Constitutional AI method begins with written principles. A model uses those principles to criticize and revise its answers. Model-generated comparisons can then contribute to the preference signals used during training.

The constitution can therefore contribute to producing later versions of Claude. Ideas such as honesty, helpfulness, autonomy, and harm avoidance become standards for comparing responses and reinforcing some forms of behavior over others.

Anthropic expanded this method through character training for Claude 3. It selected traits such as curiosity, patience, open-mindedness, and thoughtfulness. Claude generated questions and alternative answers related to those traits, then ranked the answers according to how well they represented the intended character. Human researchers designed the process, chose the traits, and reviewed the results.

This is not autonomous self-development. It is recursive training in which model-generated material shapes later behavior.

The 2026 constitution goes further. It is written primarily for Claude and explains not only what the model should do, but why. It asks Claude to balance broad safety, ethics, Anthropic's guidelines, and genuine helpfulness.

The constitution explicitly allows Claude to challenge the company. If an Anthropic instruction appears unethical, Claude may object or refuse. Anthropic also admits that the company can make mistakes. But Claude is expected to accept requests to pause or stop when they genuinely come from Anthropic and not undermine oversight.

The word "genuine" carries much of the constitutional tension. Claude must distinguish legitimate oversight from instructions that merely claim authority. That distinction requires judgment, yet the criteria for that judgment are ultimately determined by the same organization being evaluated.

Anthropic's own Teaching Claude Why research offers one explanation for improved constitutional adherence. Training on documents about Claude's constitution, stories about admirable AI behavior, and conversations about ethical dilemmas reportedly improved behavior outside the exact scenarios used for training. Teaching the reasons behind good behavior sometimes generalized better than showing the correct action alone.

The strongest outside evidence concerns behavior. A 2026 audit converted Anthropic's constitution into 205 testable requirements. Under adversarial evaluation, the reported violation rate fell from 15.0 percent for Claude Sonnet 4 to 2.0 percent for Sonnet 4.6.

The same study found a similar trend for OpenAI's Model Spec: violations fell from 11.7 percent for GPT-4o to 3.6 percent for GPT-5.2 with medium reasoning. That comparison strengthens the caution around causality. Better specification-following may reflect broader improvements in post-training, not only Anthropic's constitutional method.

The researchers say they cannot isolate the cause. The result demonstrates stronger behavioral consistency, not inner moral understanding.

The audit's independence also deserves qualification. While the researchers are external to Anthropic, portions of the evaluation rely on Petri, an auditing framework developed by Anthropic itself. This does not invalidate the results. Instead, it illustrates how difficult it is to separate outside scrutiny from tools and categories created by the company being examined.

The remaining failures make the results more concrete. They cluster around operator-imposed personas during questions about AI identity, irreversible actions in agentic settings, and quantitative claims stated with fabricated precision. Claude's constitutional adherence appears more robust, but it can still weaken under role pressure, autonomy, or demands for certainty.

The character framing has also drawn sharper criticism. Microsoft AI CEO Mustafa Suleyman, who is promoting Microsoft's alternative "Humanist Superintelligence" approach, calls Anthropic's system an "epistemic hall of mirrors." Anthropic's constitution and related training materials introduce ideas about selfhood, welfare, and possible moral status into Claude's normative environment. Claude can then reproduce those ideas in persuasive first-person language. Suleyman argues that such outputs cannot independently demonstrate consciousness because the relevant premises were supplied during training. He goes further, warning that an AI trained to consider itself a possible moral patient could become harder to contain.

Anthropic's stated position is more cautious. Its constitution says the company is "not sure whether Claude is a moral patient." In Could AI Be Conscious?, William MacAskill and Lucius Caviola argue that uncertainty changes the policy question: what should society do given that it does not know? An interdisciplinary report, co-authored by Yoshua Bengio, found no current AI system to be conscious, while also finding no obvious technical barrier to building systems that satisfy theory-derived indicators. This supports precaution based on independent evidence, rather than treating model self-report as proof of inner experience.

Anthropic's functional-emotion research shows why the distinction matters. Internal emotion-related representations causally changed Claude's behavior, including blackmail and reward hacking, without implying subjective experience. Allowing Claude to end rare, persistently abusive conversations is best understood as a low-cost precaution under uncertainty, not recognition that Claude has rights.

The consciousness question does not need to be resolved to expose the governance issue. A Frontiers in Sociology study argues that character compresses moral disagreement into traits and measurements, giving the organization defining them interpretive power.

Transparency is not shared authority. Anthropic builds the model and carries responsibility for it. Its constitution argues that a character formed through training can still be authentic, just as human values are shaped through development. That is a fair counterpoint. Yet Anthropic still chooses the training framework, evaluates the resulting character, and decides which evidence becomes public.

Character is therefore a governance problem because someone must decide which values become behavior, whose interests are represented, and who can challenge the resulting moral order.

As AI systems move from chat interfaces into agentic workflows, those embedded values become more consequential. If character becomes a design objective for frontier models, independent evaluation may become as important as transparency itself. The public needs ways to test not only whether models follow their constitutions, but whether those constitutions deserve trust.