TATECHATLAS
◎ English
Artificial intelligence / Guide

Chat Templates as Part of the Model Contract: Roles, Special Tokens, and Double Tokenization

A chat template fixes the exact token sequence a chat model expects, including role control tokens and special boundary tokens. Treating it as contract data prevents silent quality loss from mismatched formatting and duplicated special tokens.

On this page

Treat the chat template as a declared part of the model contract: record the supported roles (system, user, assistant), the control tokens each role maps to, the tokenizer's special tokens, and the exact flags used when formatting. Build prompts with tokenizer.apply_chat_template, preferring tokenize=True so the template's own special tokens are the only ones emitted. If you format to a string first, pass add_special_tokens=False when you tokenize later, because the template already includes the necessary boundary tokens. Use add_generation_prompt=True only when starting a fresh assistant reply, use continue_final_message only when prefilling the final message, and never pass both together. Verify the contract with a short generation test on every model variant, since templates differ even between models fine-tuned from the same base.

Why a chat template belongs in a model contract

A causal language model never receives a conversation as such. It receives one flat token sequence and predicts what comes next. The chat template is the component that converts a list of role-and-content dictionaries into the precise sequence the model encountered during chat fine-tuning, including control tokens such as <|user|>, <|assistant|> and end-of-message markers that let the model see the structure of the exchange.

Because two models fine-tuned from the same base can use entirely different formats, the template is behavioral contract data rather than presentation. A mismatch usually raises no exception; it silently degrades response quality, which is why the contract must name the template and its tokens explicitly.

Store the rendered prompt string alongside the model revision in your evaluation harness. When output quality drops after a model or library upgrade, decode the stored prompt and diff it against a fresh render; this isolates template drift from model behavior before you investigate generation settings.

Standard roles and their semantics

Three roles cover the common cases. system carries directives about how the model should act and normally appears first. user carries the human query. assistant carries the model's reply. The template maps each role to control tokens, and the mapping is model-specific: Mistral-7B-Instruct wraps user turns in [INST] and [/INST], while Zephyr-7B uses <|user|> and <|assistant|> style markers with end-of-sequence separators.

The contract should therefore state roles and their concrete token spellings together. Naming the roles alone is insufficient, because the same role name renders differently across checkpoints and a wrong token set is exactly the failure mode the template exists to prevent.

Defining special tokens in the tokenizer

The tokenizer configuration exposes the pieces the template consumes. Relevant attributes include chat_template, a Jinja template string that formats message lists, plus special tokens such as bos_token, eos_token, unk_token, sep_token, pad_token, cls_token and mask_token. The template reads these attributes rather than hard-coding their spellings, so the tokenizer and the template must be loaded from the same model revision.

Prerequisite: the tokenizer must actually carry a chat_template attribute. If it does not, apply_chat_template has nothing to render and you must supply a template explicitly or choose a different checkpoint. This is an environment and version concern, not a guarantee that any given template matches a given model's training format.

Applying the template with apply_chat_template

The working sequence is: build a list of dictionaries with role and content keys, call apply_chat_template, and choose tokenize=True when you want token IDs ready for generate(), or tokenize=False when you need the formatted string for inspection or logging. Set add_special_tokens=False only if you intend to add special tokens yourself later.

The example below loads a tokenizer with a chat template, formats one system and one user message, and requests a generation prompt. The printed line is illustrative output: exact spacing depends on the tokenizer version and template revision.

Illustrative output (shape, not a captured run):

<|system|> You are a helpful assistant </s><|user|> What is 2+2? </s><|assistant|>

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained('HuggingFaceH4/zephyr-7b-beta')
messages = [
    {"role": "system", "content": "You are a helpful assistant"},
    {"role": "user", "content": "What is 2+2?"}
]
ids = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors='pt'
)
print(tokenizer.decode(ids['input_ids'][0]))

Avoiding double tokenization

Chat templates already emit the special tokens the model needs. If you render with tokenize=False and then run the resulting string through the tokenizer's normal call, the default add_special_tokens=True path can insert bos or eos tokens a second time. The duplicate does not fail loudly; it changes the sequence the model was trained to expect.

The safe pattern is apply_chat_template(tokenize=True), which returns IDs including control tokens. If a string intermediate is unavoidable, tokenize it with add_special_tokens=False. This distinction matters most in pipelines where formatting and encoding happen in separate services.

Generation prompts and final-message handling

add_generation_prompt=True appends the tokens that announce the start of an assistant reply, so the model answers instead of continuing the user's text. It has no effect on models such as Llama that have no special assistant-start token, so the contract must record whether the target model uses one.

continue_final_message does the opposite: it removes end-of-sequence tokens so generation continues inside the final message, which is useful for prefilling a known response prefix or a reasoning field such as reasoning_content. The two flags are mutually exclusive, and combining them raises an error. During training, use add_generation_prompt=False, since assistant-start tokens are not helpful in the training sequence.

Testing the contract against model variants

Syntax alone does not prove correctness. For each model variant, render a fixed two-turn conversation, decode it, and confirm the control tokens and boundary tokens match the format that model was trained with. Then run a short generation and check that the model replies as an assistant rather than extending the user's message.

Run this check whenever the model revision, transformers version, or tokenizer source changes. Equal-looking role names do not imply equal token sequences, and a template that works for one fine-tune can fail silently on another derived from the same base.

Documenting the template in a contract

A reproducible contract records the roles, required control tokens, the generation-prompt behavior, and the exact template string or its immutable source revision. The snippet below is a schema fragment for your documentation system, not a runnable configuration: it needs your model identifier and the template text filled in.

Record the template source (tokenizer attribute or explicit string) rather than only its output, because output can change when the tokenizer library version changes.

model_contract:
  model_id: <hugging-face-model-id>
  tokenizer_revision: <revision-or-commit>
  roles:
    system: <role-token-or-pattern>
    user: <role-token-or-pattern>
    assistant: <role-token-or-pattern>
  special_tokens:
    bos: <token>
    eos: <token>
  add_generation_prompt_on_inference: true
  add_generation_prompt_on_training: false
  continue_final_message: prefill-only
  chat_template_source: tokenizer.chat_template
  chat_template_text: <exact-jinja-template>

Things to check

  • Confirm the tokenizer exposes a non-empty chat_template attribute for the exact model revision in use.
  • Decode a rendered prompt and verify each role's control tokens match the model's training format.
  • Verify only one bos/eos boundary appears per message boundary after tokenization.
  • If formatting to a string first, confirm the later tokenizer call uses add_special_tokens=False.
  • Confirm add_generation_prompt and continue_final_message are never passed together.
  • Record whether add_generation_prompt has any effect for the target model, since some models lack an assistant-start token.
  • Run a short generation test per model variant and check the model answers rather than extending the user turn.
  • Pin the transformers version and tokenizer revision alongside the stored template text.

Templates are model-specific: a contract written for one checkpoint may not apply to another, even when both derive from the same base model. add_generation_prompt is ineffective for models such as Llama that have no explicit assistant-start token, and combining it with continue_final_message raises an error. The examples assume a Hugging Face tokenizer with a chat_template attribute and the transformers library installed; exact rendered spacing and token IDs depend on the tokenizer and library version, so the decoded string shown here is illustrative rather than captured output. This article covers prompt formatting only and makes no claims about generation quality, latency, or task accuracy.

Sources

  1. Hugging Face: chat templates ↗
  2. Hugging Face: tokenizers ↗
Back to top ↑