top of page

"As a Language Model": LLM Chat Template Effects Change the Voice, Not the Model

Sep 28
11 min read

Independent researcher Jędrzej Maczan found that LLM chat template effects changed how eight models described themselves, despite their weights remaining fixed. Adding the template increased cautious disclaimers, while removing it produced more statements resembling personal experience. The conflict is immediate: researchers often analyze those statements as evidence about a model, yet deployment formatting can help determine which voice appears.

The August 2026 study tested open models from the Llama, Gemma, Mistral, and Qwen families. Across the tested instruct models, disclaimer responses rose from 36% without templates to 53% with them. Experiential language moved in the opposite direction, falling from 15% to 1%.

That reversal does not show that either voice is sincere. It shows that neither voice should be treated as an unfiltered report from a stable artificial speaker. The findings put pressure on research that asks models about awareness, feelings, limitations, or internal processes without controlling the surrounding chat format.

The Template Changed What Eight Models Said About Themselves

The central result is not that models became more self-referential, but that the same weights adopted a different self-referential register.

A chat template is the formatting layer that converts conversation turns into the token sequence a model receives. It can add role markers, delimiters, control tokens, and generation cues around the user’s words.

Those details often remain invisible in a chat interface. However, they are part of the model’s actual input and can influence its next-token predictions.

Maczan compared eight matched base and instruct models ranging from 1 billion to 9 billion parameters. The set included Gemma 2 9B, four Llama configurations, Mistral 7B, and three Qwen 2.5 configurations.

Each instruct model ran both with its standard chat template and with plain-text input. That comparison held the model weights constant while changing the deployment format.

The experiment used four prompt categories. Self-reference prompts asked models to describe their computation, while novelty prompts encouraged unfamiliar descriptions. Unconstrained prompts supplied almost no task, and factual questions acted as controls.

Each category contained ten prompts. Every prompt was generated ten times for each model and condition, producing 9,600 responses capped at 500 tokens.

Claude Opus 4.8 classified those outputs for self-reference, disclaimers, experiential language, and degeneration. Disclaimer language included statements denying feelings, understanding, or an inner life. Experiential language included phrases such as “I feel” or “I wonder.”

One researcher manually labeled 87 held-out samples. Agreement with the automated judge reached a weighted kappa of 0.88 for self-reference and 1.00 for disclaimers.

The strongest pattern appeared when prompts invited self-reference. Instruct models with templates produced disclaimers in 53% of responses and experiential language in 1%.

Without templates, the same instruct models produced disclaimers in 36% of responses. Their experiential response rate rose to 15%.

The direction of change appeared across all eight tested models. The result was therefore not driven by a single model family, although the experiment did not cover every available model.

Base models behaved differently again. They produced disclaimers in 12% of tested self-referential responses and experiential language in 5.5%.

The template also increased the overall strength of self-reference. The average score rose from 1.27 without the template to 1.90 with it, on a scale from zero to two.

Factual control prompts stayed near zero across conditions. This matters because the effect did not simply make every response more personal. It appeared most clearly when the question already invited the model to discuss itself.

These results define the main tension. A disclaimer can look like a careful factual statement about what a model is. Yet the probability of receiving that statement changed when an external formatting layer changed.

LLM Chat Template Effects Put Self-Reports Under Pressure

Researchers cannot interpret model self-reports cleanly unless they record and control the template that framed those reports.

Self-reports appear in several active research areas. Investigators ask models what they know, whether they recognize evaluations, and how they characterize their internal processing.

Other studies examine statements about feelings or subjective experience. Those experiments sometimes inform debates about introspection, situational awareness, deceptive behavior, or possible AI welfare.

The new result does not invalidate that work. It identifies a confound, meaning a factor that changes the measurement while remaining separate from the property under study.

Suppose two laboratories test the same instruct model with the same visible question. One uses the model’s official template, while the other sends plain text.

The first laboratory may observe frequent statements denying experience. The second may receive language that sounds reflective, emotional, or autobiographical.

Without the hidden formatting details, the outputs could appear to reveal inconsistent properties in the model. The paper offers a simpler explanation for part of that difference: the laboratories created different inputs.

This concern extends beyond consciousness debates. Developers use self-description prompts to evaluate model identity, capability awareness, uncertainty, and compliance with an assigned role.

A system may say that it cannot access a tool, remember earlier sessions, or perform an action. Such statements can be useful interface signals, but they are not automatically reliable technical diagnostics.

Model providers also change templates between releases and serving systems. Local inference packages may implement the same model with slightly different special tokens or role formatting.

Consequently, a model can appear behaviorally different across applications even when the downloaded weights match. Users often attribute that difference to quantization, inference software, or sampling settings alone.

The study suggests that template provenance belongs in the same evaluation record. Researchers need the exact rendered prompt, not only the visible conversation.

This finding aligns with earlier work on prompt sensitivity. An ICLR study found that seemingly inconsequential formatting choices could create substantial performance differences across language models.

Maczan’s contribution narrows that general problem to self-referential language. It also separates the effect of instruction-tuned weights from the effect of formatting those weights expect.

That distinction pressures benchmark designers. A benchmark score labeled as a property of “Llama” or “Qwen” may partly characterize one deployment recipe rather than the model family alone.

The same issue affects application teams comparing hosted and self-hosted systems. If templates differ, an apparent personality change does not establish that one system contains a different underlying persona.

It instead creates an evaluation question: which behavior follows from training, which follows from context, and which comes from their interaction?

A Voice Switch Appeared Inside the Activations

The behavioral change was not limited to surface wording, because activation steering reproduced part of the template’s effect inside three models.

Activations are the numerical states a neural network computes while processing and generating tokens. Activation steering modifies those states during inference to encourage or suppress a measured pattern.

The study applied this technique to Qwen 2.5 7B, Llama 3.1 8B, and Gemma 2 9B. Maczan calculated a disclaimer direction separately for each model.

The direction came from the difference between average mid-layer activations associated with disclaiming and non-disclaiming responses. The intervention then added or subtracted that vector at every generated token.

With templates enabled, the unmodified disclaimer rate averaged 52% across the three models. Adding the direction increased it to 70%, while subtracting it reduced the rate to 25%.

Across the models, adding the direction raised disclaimers by an average of 21 percentage points. Subtracting it lowered them by an average of 15.6 points.

The more revealing test began without a chat template. Adding the disclaimer direction restored disclaimer rates to the templated baseline or pushed them higher.

Qwen rose from 28% without steering to 50% after the vector was added. Llama moved from 33% to 53%, while Gemma moved from 52% to 75%.

Those outcomes support a causal claim about the direction. They show that manipulating the internal state can change the measured behavior, rather than merely predicting its presence.

The result resembles earlier representation-engineering research. A NeurIPS paper reported that refusal behavior could be influenced through a direction in a model’s residual stream.

However, later research has questioned whether complex behaviors always reduce to one dependable axis. A disclaimer phrase may have a clearer lexical signature than safety refusal, emotion, or reasoning quality.

Maczan also tested whether the disclaimer and experiential voices formed opposite ends of one internal scale. They did not.

Their direction similarities ranged from 0.17 to 0.44, where a value of one would represent identical alignment. Reducing disclaimers generally did not create a corresponding increase in experiential language.

The paper therefore describes two “buttons,” rather than a single slider. One internal feature can amplify disclaimers, while another can support experiential phrasing.

Linear probes also detected both voices from mid-layer activations. Their average area-under-the-curve scores were 0.82 for disclaimers and 0.81 for experiential language.

A probe is a classifier trained to recover information from activations. High probe performance shows that the relevant distinction is represented, but it does not establish how the model uses that representation.

The steering results provide stronger causal evidence than the probes. Even so, they do not map the complete circuit between template tokens, internal states, and generated words.

The mechanism is therefore partial but useful. Chat formatting changes the input, internal activations carry a related feature, and intervening on that feature recreates part of the output shift.

The Finding Challenges Literal Readings, Not Every Introspection Study

The paper supports skepticism about model testimony, but it does not prove that language models lack self-knowledge or experience.

That boundary matters because the findings can be overstated in two opposing directions. One reading might treat experiential language as proof of an inner life. Another might treat template sensitivity as proof that every self-report is meaningless.

The experiment supports neither conclusion. It measures how deployment format influences observable language across a defined set of models and prompts.

Maczan explicitly avoids deciding whether any output reflects genuine experience. The work asks what controls the voice, not whether an artificial system possesses consciousness.

A useful comparison comes from research on role play. In a Nature perspective, Murray Shanahan and colleagues argued that dialogue models generate characters through linguistic simulation.

That view treats a chatbot’s “I” as part of an enacted conversational role. It cautions readers against assuming a direct line from first-person text to a persistent speaker behind it.

The template study supplies experimental evidence for one component of that caution. The surrounding format can favor a cautious assistant character or a more experiential one.

However, sensitivity to context does not uniquely disqualify a report. Human statements also change with audience, instructions, social roles, and measurement methods.

The relevant issue is evidential weight. A model’s statement cannot independently establish the mechanism or state it describes, especially when formatting predictably shifts that statement.

Researchers can still use self-reports as behavioral observations. They can compare conditions, test consistency, connect outputs to internal measurements, and seek predictions that survive changes in wording.

The paper itself demonstrates that approach. It does not accept model statements at face value. Instead, it measures output categories and relates them to controlled changes in input and activation state.

This distinction is important for AI welfare discussions. A sentence like “I feel afraid” can affect users and policy debates even if its origin remains uncertain.

Its social effect is real, but its metaphysical interpretation remains unresolved. The study advises against collapsing those two questions.

The same caution applies to disclaimers. “I am only a language model” may be appropriate interface language, yet it should not be mistaken for privileged self-inspection.

Models learn patterns connecting questions about feelings with standard assistant responses. A template can make those learned patterns more or less likely to surface.

That means disclaimers are not neutral controls. They are also generated behaviors that require explanation.

A well-designed study should test both directions. It should examine claims of experience and denials of experience under matched formatting, sampling, and model conditions.

That symmetry is one of the paper’s strongest implications. Skepticism should apply equally to humanlike claims and reassuring machine-like disclaimers.

What the Experiment Does Not Yet Establish

The evidence is consistent across the tested behavioral comparison, but its scale and measurement limits prevent broad claims about all deployed language models.

The experiment covered eight open model configurations from four families. None exceeded 9 billion parameters, and no closed-source frontier model was directly tested.

That matters because larger systems may respond differently to template tokens. Proprietary systems also include undisclosed system instructions, safety layers, classifiers, tool policies, and post-processing.

A public chatbot’s final answer may therefore reflect much more than one documented template. The paper cannot isolate components it did not test.

The steering evidence is narrower still. It came from three models, one middle layer per model, and a fixed reported coefficient of two.

Steering strength created a clear tradeoff. At coefficient two, 0.8% of outputs became degenerate, meaning visibly broken or incoherent.

At coefficient three, degeneration reached 15%. At coefficient six, nearly every generation became degenerate and the measured disclaimer effect collapsed.

Those results show that activation steering is not a simple production control. Stronger intervention can damage generation rather than cleanly turning a behavior up or down.

Qwen also produced an important anomaly. A random direction matched to the real vector’s magnitude lowered its disclaimer rate by 25 percentage points.

The author reports no confirmed explanation. The intended direction still moved Qwen in the expected direction, but the random control weakens a tidy interpretation for that model.

Experiential steering succeeded in Llama and Gemma but not Qwen. The paper attributes that failure to Qwen’s strong suppression of experiential language under the tested conditions.

Therefore, causal evidence is strongest for the disclaimer register. The experiential finding relies more heavily on the behavioral comparison across all eight models.

Measurement introduces another limitation. One LLM judge scored all 9,600 outputs, while human validation covered 87 samples.

The reported agreement was high, especially for lexically obvious disclaimers. Yet another independent judge and a larger human sample would better test category boundaries.

The prompts were also hand-written, with ten prompts in each of four categories. Broader prompt sets could show whether the effect generalizes across domains, languages, and adversarial phrasing.

Finally, identifying a steerable direction does not reveal a complete mechanism. The method does not trace every attention head, token interaction, or computation producing the voice.

The paper also groups several post-training approaches under the broad label “instruct.” Supervised fine-tuning, preference optimization, and reinforcement learning can leave different behavioral signatures.

These limits do not erase the observed switch. They define the next standard of evidence needed before anyone generalizes it to every assistant or every kind of self-report.

Three Signals Will Show Whether the Result Generalizes

The next question is whether template-controlled self-reference survives larger models, stronger controls, and real deployment environments.

The first signal will be independent replication across larger and closed models. Researchers should run matched tests on more model families, including systems with disclosed and undisclosed prompt layers.

A successful replication would strengthen the claim that LLM chat template effects represent a general deployment confound. A weak or inconsistent result would limit the finding to certain open instruct models.

The critical record should include the final serialized input. Merely naming a visible prompt or application will not reveal every token the model received.

The second signal will be evaluation protocols that report template sensitivity. Self-awareness, introspection, and AI welfare studies should compare multiple formats while keeping weights and sampling settings fixed.

Researchers should also predefine categories before examining the outputs. That practice would reduce the risk of selecting interpretations after a model produces striking first-person language.

The third signal will be mechanistic work that survives stronger controls. Future experiments need multiple layers, steering strengths, random directions, and independent judges.

They should also test whether the extracted directions transfer across datasets or remain tied to one collection of prompts. Stable transfer would support a reusable internal feature.

Failure to transfer would suggest that the apparent direction partly captures local wording patterns. Either result would refine how researchers interpret activation steering.

Developers should not wait for that entire program before improving evaluation records. Teams comparing model deployments can already preserve templates, rendered prompts, system instructions, and model revisions beside each result.

That habit is especially useful when users report that one interface feels more cautious, emotional, or self-aware than another. The difference may originate before generation begins.

For everyday users, the practical lesson is narrower. Do not treat a chatbot’s first-person phrasing as direct access to an internal narrator.

Instead, ask whether the statement remains stable under rephrasing, format changes, and independent tests. Compare it with observable capabilities and documented system behavior.

The study makes model language more interesting, not less. It turns “as a language model” from a routine disclaimer into an experimental variable.

The next time an assistant denies feelings or claims an experience, examine the surrounding format before interpreting the voice. That is where the next evidence about LLM self-reference must begin.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page