Welcome to spark

Your guide to understanding the codebase

Learning Path

16 units • 5 levels

Orchestration & APIs

  • Pipelines, workflows, public interfaces • 3 units

Spark Core Architecture

  • Architecture · 19

Core Data Abstractions

  • Data Model · 5

Configuration and Infrastructure Basics

  • Infrastructure · 5

Core Logic & Data

  • Business rules, schemas, models • 3 units

Communication and RPC

  • Infrastructure · 18

Monitoring and Observability

  • Infrastructure · 10

Distributed Execution Primitives

  • Workflow · 6

Interaction & Integration

  • UI components, external connectors • 3 units

APIs and Interfaces

  • API · 14

Security Infrastructure

  • Security · 2

Resource and Task Management

  • Workflow · 7

Cross-Cutting Concerns

  • Auth, logging, config, testing • 2 units

External Data Systems

  • Integration · 20

Language Integrations

  • Integration · 18

Edge Cases & Resilience

  • Error handling, fault tolerance • 2 units

Additional Data Model Patterns

  • Data Model · 24

Additional Error Handling Patterns

  • Error Handling · 3

Test Your Knowledge

Test your deep understanding of the codebase

Progress: 0/26 answered

Question Tiers - ordered for deep understanding

  • Why: 8
  • Purpose & Problem: Architecture: 10
  • Design & Patterns: Code: 8

Hands-On Assignment

45-75 minutes

Your Challenge

Extend Spark's error handling system to support error severity levels (INFO, WARNING, ERROR, CRITICAL) and implement a custom InMemoryStore-backed error event tracker that records recent errors with their severity. Your implementation should allow querying errors by severity level and provide statistics on error frequency, integrating seamlessly with the existing ErrorClassesJSONReader infrastructure.

Starting Points

  • common/utils/src/main/resources/error/error-classes.json:1-86
    Examine the current error class structure and JSON schema
    explore
  • common/utils/src/main/scala/org/apache/spark/ErrorClassesJSONReader.scala:1-83
    Study how error classes are loaded and parsed from JSON
    explore
  • common/kvstore/src/main/java/org/apache/spark/util/kvstore/InMemoryStore.java:1-113
    Understand the KVStore interface and InMemoryStore implementation - this will be your storage backend
    reference
  • common/utils/src/main/resources/error/error-classes.json
    Add severity field to error class definitions here
    modify
  • common/utils/src/main/scala/org/apache/spark/ErrorClassesJSONReader.scala
    Extend the reader to parse severity levels from error definitions
    modify

Success Criteria

  • At least 5 existing error classes in error-classes.json have severity levels assigned
  • ErrorClassesJSONReader successfully parses severity field without breaking existing functionality
  • ErrorEventTracker can store and retrieve at least 100 error events using InMemoryStore
  • Can query errors by severity level (e.g., getAllCriticalErrors()) and get accurate results
  • Statistics method returns correct counts per severity level over a time window
  • Existing Spark error handling tests still pass without modification

Hints

  1. Understanding the error class architecture conceptual
  2. InMemoryStore usage pattern code location
  3. Integrating with existing JSON schema implementation
  4. Building the tracker component implementation

Prerequisites

  • Scala case classes
  • JSON parsing
  • Key-value store concepts
  • Java generics