Welcome to transformers

Your guide to understanding the codebase


Learning Path

5 levels
15 learning units

Orchestration & APIs

  • Pipelines, workflows, public interfaces • 3 units
    • Pipeline Abstraction and Basic Usage
    • Model Configuration and Auto Classes
    • Tokenization and Data Processing

Core Logic & Data

  • Business rules, schemas, models • 3 units
    • Transformer Architecture Fundamentals
    • Feature Extraction and Generation
    • Training Framework and Callbacks

Interaction & Integration

  • UI components, external connectors • 3 units
    • Multimodal and Specialized Architectures
    • Distributed Training and Model Sharding
    • Model Hub and Integration Ecosystem

Cross-Cutting Concerns

  • Auth, logging, config, testing • 2 units
    • Model Optimization and Quantization
    • Testing, Benchmarking, and Deployment

Edge Cases & Resilience

  • Error handling, fault tolerance • 1 unit
    • Additional Data Model Patterns

Test Your Knowledge

24 questions across all levels

Start Quiz

Progress
0/24 answered

Hands-On Assignment

2-3 hours

Your Challenge

Implement a custom tokenizer that supports a new special token type called 'ENTITY' tokens for named entity recognition tasks. Your tokenizer should extend the existing PreTrainedTokenizer class, handle entity markers (e.g., <ENTITY:PERSON>, <ENTITY:ORG>) during encoding/decoding, and properly integrate with the AutoTokenizer registry so it can be loaded via AutoTokenizer.from_pretrained(). The implementation should maintain compatibility with existing tokenization pipelines while adding entity-aware vocabulary management.

Starting Points

  • src/transformers/tokenization_utils_base.py
    Explore the PreTrainedTokenizerBase class to understand the core tokenization interface, special token handling, and encoding/decoding methods

  • src/transformers/tokenization_utils.py
    Study PreTrainedTokenizer implementation - this is the class you'll extend for your custom tokenizer

  • src/transformers/models/bert/tokenization_bert.py
    Reference this as a concrete example of how tokenizers are implemented, including __init__, vocab loading, and special methods

  • src/transformers/models/auto/tokenization_auto.py
    Understand how AutoTokenizer registry works - you'll need to register your custom tokenizer here

  • src/transformers/models/
    Create a new directory here for your custom tokenizer implementation (e.g., 'entity_aware/')

  • tests/models/bert/test_tokenization_bert.py
    Review tokenizer testing patterns - you'll create similar tests for your implementation

Success Criteria

  • Your tokenizer can encode text with entity markers like 'John ENTITY:PERSON works at Google ENTITY:ORG' and preserve entity information
  • Decoding returns the original text with entity markers intact
  • The tokenizer can be saved with save_pretrained() and loaded with AutoTokenizer.from_pretrained()
  • Entity type vocabulary persists across save/load cycles
  • At least 5 unit tests pass covering: basic tokenization, special token handling, save/load, entity type management, and edge cases
  • The tokenizer works with a simple text classification pipeline without errors

Hints

  • (click to reveal)
    • Hint 1: Understanding special tokens architecture conceptual
    • Hint 2: Key methods to override code location
    • Hint 3: Registry integration pattern implementation
    • Hint 4: Handling dynamic entity types implementation
    • Hint 5: Testing your implementation implementation

Prerequisites

  • Python class inheritance
  • Tokenization concepts (vocab, encoding, decoding)
  • JSON serialization
  • Basic understanding of transformer model inputs