# Welcome to transformers

Your guide to understanding the codebase

---

## Learning Path

**5 levels**  
**15 learning units**

### Orchestration & APIs  
- Pipelines, workflows, public interfaces • 3 units  
  - Pipeline Abstraction and Basic Usage  
  - Model Configuration and Auto Classes  
  - Tokenization and Data Processing

### Core Logic & Data  
- Business rules, schemas, models • 3 units  
  - Transformer Architecture Fundamentals  
  - Feature Extraction and Generation  
  - Training Framework and Callbacks

### Interaction & Integration  
- UI components, external connectors • 3 units  
  - Multimodal and Specialized Architectures  
  - Distributed Training and Model Sharding  
  - Model Hub and Integration Ecosystem

### Cross-Cutting Concerns  
- Auth, logging, config, testing • 2 units  
  - Model Optimization and Quantization  
  - Testing, Benchmarking, and Deployment

### Edge Cases & Resilience  
- Error handling, fault tolerance • 1 unit  
  - Additional Data Model Patterns

## Test Your Knowledge  
**24 questions across all levels**  
### Start Quiz  
**Progress**  
0/24 answered

## Hands-On Assignment  
**2-3 hours**  
### Your Challenge  
Implement a custom tokenizer that supports a new special token type called 'ENTITY' tokens for named entity recognition tasks. Your tokenizer should extend the existing PreTrainedTokenizer class, handle entity markers (e.g., `<ENTITY:PERSON>`, `<ENTITY:ORG>`) during encoding/decoding, and properly integrate with the AutoTokenizer registry so it can be loaded via AutoTokenizer.from_pretrained(). The implementation should maintain compatibility with existing tokenization pipelines while adding entity-aware vocabulary management.

### Starting Points  
- `src/transformers/tokenization_utils_base.py`  
  Explore the PreTrainedTokenizerBase class to understand the core tokenization interface, special token handling, and encoding/decoding methods
  
- `src/transformers/tokenization_utils.py`  
  Study PreTrainedTokenizer implementation - this is the class you'll extend for your custom tokenizer
  
- `src/transformers/models/bert/tokenization_bert.py`  
  Reference this as a concrete example of how tokenizers are implemented, including `__init__`, vocab loading, and special methods  
  
- `src/transformers/models/auto/tokenization_auto.py`  
  Understand how AutoTokenizer registry works - you'll need to register your custom tokenizer here
  
- `src/transformers/models/`  
  Create a new directory here for your custom tokenizer implementation (e.g., 'entity_aware/')
  
- `tests/models/bert/test_tokenization_bert.py`  
  Review tokenizer testing patterns - you'll create similar tests for your implementation

### Success Criteria  
- Your tokenizer can encode text with entity markers like 'John <ENTITY:PERSON> works at Google <ENTITY:ORG>' and preserve entity information  
- Decoding returns the original text with entity markers intact  
- The tokenizer can be saved with `save_pretrained()` and loaded with `AutoTokenizer.from_pretrained()`  
- Entity type vocabulary persists across save/load cycles  
- At least 5 unit tests pass covering: basic tokenization, special token handling, save/load, entity type management, and edge cases  
- The tokenizer works with a simple text classification pipeline without errors

### Hints  
- (click to reveal)  
  - Hint 1: Understanding special tokens architecture conceptual  
  - Hint 2: Key methods to override code location  
  - Hint 3: Registry integration pattern implementation  
  - Hint 4: Handling dynamic entity types implementation  
  - Hint 5: Testing your implementation implementation

### Prerequisites  
- Python class inheritance  
- Tokenization concepts (vocab, encoding, decoding)  
- JSON serialization  
- Basic understanding of transformer model inputs
