Welcome to transformers
Your guide to understanding the codebase
Learning Path
5 levels
15 learning units
Orchestration & APIs
- Pipelines, workflows, public interfaces • 3 units
- Pipeline Abstraction and Basic Usage
- Model Configuration and Auto Classes
- Tokenization and Data Processing
Core Logic & Data
- Business rules, schemas, models • 3 units
- Transformer Architecture Fundamentals
- Feature Extraction and Generation
- Training Framework and Callbacks
Interaction & Integration
- UI components, external connectors • 3 units
- Multimodal and Specialized Architectures
- Distributed Training and Model Sharding
- Model Hub and Integration Ecosystem
Cross-Cutting Concerns
- Auth, logging, config, testing • 2 units
- Model Optimization and Quantization
- Testing, Benchmarking, and Deployment
Edge Cases & Resilience
- Error handling, fault tolerance • 1 unit
- Additional Data Model Patterns
Test Your Knowledge
24 questions across all levels
Start Quiz
Progress
0/24 answered
Hands-On Assignment
2-3 hours
Your Challenge
Implement a custom tokenizer that supports a new special token type called 'ENTITY' tokens for named entity recognition tasks. Your tokenizer should extend the existing PreTrainedTokenizer class, handle entity markers (e.g., <ENTITY:PERSON>, <ENTITY:ORG>) during encoding/decoding, and properly integrate with the AutoTokenizer registry so it can be loaded via AutoTokenizer.from_pretrained(). The implementation should maintain compatibility with existing tokenization pipelines while adding entity-aware vocabulary management.
Starting Points
src/transformers/tokenization_utils_base.py
Explore the PreTrainedTokenizerBase class to understand the core tokenization interface, special token handling, and encoding/decoding methodssrc/transformers/tokenization_utils.py
Study PreTrainedTokenizer implementation - this is the class you'll extend for your custom tokenizersrc/transformers/models/bert/tokenization_bert.py
Reference this as a concrete example of how tokenizers are implemented, including__init__, vocab loading, and special methodssrc/transformers/models/auto/tokenization_auto.py
Understand how AutoTokenizer registry works - you'll need to register your custom tokenizer heresrc/transformers/models/
Create a new directory here for your custom tokenizer implementation (e.g., 'entity_aware/')tests/models/bert/test_tokenization_bert.py
Review tokenizer testing patterns - you'll create similar tests for your implementation
Success Criteria
- Your tokenizer can encode text with entity markers like 'John ENTITY:PERSON works at Google ENTITY:ORG' and preserve entity information
- Decoding returns the original text with entity markers intact
- The tokenizer can be saved with
save_pretrained()and loaded withAutoTokenizer.from_pretrained() - Entity type vocabulary persists across save/load cycles
- At least 5 unit tests pass covering: basic tokenization, special token handling, save/load, entity type management, and edge cases
- The tokenizer works with a simple text classification pipeline without errors
Hints
- (click to reveal)
- Hint 1: Understanding special tokens architecture conceptual
- Hint 2: Key methods to override code location
- Hint 3: Registry integration pattern implementation
- Hint 4: Handling dynamic entity types implementation
- Hint 5: Testing your implementation implementation
Prerequisites
- Python class inheritance
- Tokenization concepts (vocab, encoding, decoding)
- JSON serialization
- Basic understanding of transformer model inputs