Python Transformers: Padding, Truncation, and Attention Masks
Understand how Hugging Face tokenizers use padding, truncation, and attention masks to shape batches for transformer models, with practical examples and common edge cases.
When you batch text for a Hugging Face transformer model, every sequence in the batch must have the same number of tokens. The tokenizer normalizes inputs with three related settings: padding, truncation, and attention masks. Padding adds special tokens to shorter sequences, truncation removes tokens from longer ones, and attention masks tell the model which positions are real content. Together they create a rectangular input tensor without letting artificial padding tokens affect the model's predictions.
Why Padding, Truncation, and Attention Masks Are Connected
Transformer models operate on fixed-shape tensors. A batch of sequences must be a single rectangular matrix, so every sequence must have the same number of tokens. Because natural text rarely comes in equal lengths, the tokenizer has to normalize the input.
Padding appends special tokens to shorter sequences so they match the length of the longest sequence in the batch. Truncation removes tokens from sequences that exceed a chosen maximum length, typically the model's maximum input length. Attention masks mark which positions are real tokens and which are padding tokens.
These three settings are not independent. Padding without truncation can let very long sequences exceed the model's limit. Truncation without padding leaves sequences at different lengths. And padding without an attention mask makes the model treat padding tokens as real content, which distorts attention weights and produces incorrect outputs.
How the Tokenizer Handles Padding and Truncation
The AutoTokenizer class exposes padding and truncation as arguments to its __call__ method. You pass padding and truncation along with max_length when tokenizing a batch of texts.
from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained('bert-base-uncased') texts = [ 'The quick brown fox jumps over the lazy dog.', 'A much shorter sentence.', ] encoded = tokenizer( texts, padding=True, truncation=True, max_length=128, return_tensors='pt', )
padding=True (or padding='longest') pads each batch to the longest sequence in that batch. truncation=True cuts sequences longer than max_length down to that limit. If you omit max_length, truncation uses tokenizer.model_max_length. Set max_length to the model's maximum if you want to allow the full context, or to a smaller value to save compute.
The resulting encoded object contains input_ids and attention_mask; for models like BERT that use token type embeddings, it also contains token_type_ids.
How Attention Masks Work
The attention mask is a tensor with the same shape as input_ids. It contains 1s for real tokens and 0s for padding tokens. The model uses it during the self-attention computation to keep padding positions from contributing to attention scores.
print(encoded['attention_mask'])
In the example above, the first sequence is longer than the second, so the second sequence receives padding tokens. Its attention mask has 1s for real tokens and 0s for padding tokens. The model's attention layers use the mask to ignore padding positions when computing attention weights, so positions with a mask value of 0 do not affect the weighted sum over the sequence.
The mask is therefore not optional. Without it, the model would attend to padding tokens, shifting the hidden representations of real tokens and degrading the output.
Choosing a Padding Mode
The padding argument accepts several values:
| Value | Behavior |
|---|---|
True or 'longest' | Pads to the longest sequence in the batch |
'max_length' | Pads every sequence to max_length tokens |
False or 'do_not_pad' | No padding |
Use padding='longest' when sequence lengths vary within a batch and you want to minimize wasted computation. Use padding='max_length' when you need a fixed tensor shape, which is common in serving pipelines that preallocate memory or expect a consistent input shape.
Dynamic padding with 'longest' is usually more efficient for training because it avoids padding short sequences up to a large fixed length. The tradeoff is that the batch tensor shape varies between batches, which is fine for most training loops but can complicate inference pipelines that expect a fixed shape.
Choosing a Truncation Strategy
The truncation argument accepts several values:
| Value | Behavior |
|---|---|
True or 'longest_first' | Truncates to max_length (or the model maximum), removing tokens from the end; for pairs, it removes from the longer sequence first |
'only_first' | Truncates only the first sequence in a pair |
'only_second' | Truncates only the second sequence in a pair |
False or 'do_not_truncate' | No truncation |
For single-sequence inputs, truncation=True is equivalent to 'longest_first': it removes tokens from the end of the sequence. For paired inputs, such as question-answer pairs or sentence pairs, 'only_first' and 'only_second' let you choose which sequence gets cut. This matters when one side of the pair is more important than the other.
encoded = tokenizer( question, context, padding=True, truncation='only_second', max_length=384, )
This is a common pattern in extractive question answering, where the context is truncated to fit inside the model's limit while the question is kept as complete as possible.
Dynamic Padding with DataCollatorWithPadding
In a training loop, batches can have varying lengths. DataCollatorWithPadding dynamically pads each batch to the longest sequence in that batch, avoiding unnecessary fixed-length padding.
from transformers import DataCollatorWithPadding data_collator = DataCollatorWithPadding(tokenizer=tokenizer, return_tensors='pt')
When you pass this collator to a Trainer or use it in a PyTorch DataLoader, it pads each batch to the batch's longest sequence and generates the corresponding attention mask automatically. This is the recommended approach for training because it keeps the batch tensor as small as possible while still meeting the model's fixed-shape requirement.
Common Mistakes and Edge Cases
One common mistake is forgetting that some tokenizers do not define a padding token. If you try to pad without a pad_token, the tokenizer raises an error. You can set one explicitly:
if tokenizer.pad_token is None: tokenizer.pad_token = tokenizer.eos_token
For decoder-only models used for generation, set tokenizer.padding_side = 'left' so padding appears before the real tokens. If padding is on the right, the last position in the sequence can be a padding token, and autoregressive generation starts from that padding position instead of the final real token.
Another edge case: setting both padding='max_length' and truncation=True does not mean every sequence is truncated. Truncation only removes tokens from sequences longer than max_length; shorter sequences are padded up to max_length, not cut down.
Performance Considerations
Padding wastes computation because the model still processes padding tokens even though the attention mask neutralizes their contribution. The amount of wasted work depends on how much padding you add. Dynamic padding with 'longest' minimizes waste within a batch, but across batches it depends on the variance of sequence lengths in your dataset.
For inference, you can sort requests by length and batch similar-length sequences together to reduce padding overhead. This is a common serving optimization, though it adds latency complexity.
The attention mask itself adds a small amount of memory per token, but it is negligible compared with model weights and activations. The real cost is the extra forward-pass work on padding tokens, which is why dynamic padding and length-based batching matter in production.