Runtime Topology
Process Descriptions
API Server: Receives external requests via REST and CLI gateway. Validates authentication, enforces rate limits, and submits jobs to the scheduler. Stateless. Horizontally scalable. Scheduler: Maintains job queues, assigns priorities, and dispatches to workers. Tracks GPU allocation state. Supports preemption for high-priority experiments. GPU Workers: Execute TensorEngine, MergeEngine, and EvolutionEngine operations. One worker owns one or more GPUs. Runs in a container with NVIDIA/ROCm drivers mounted. CPU Workers: Handle data preprocessing, tokenizer operations, and lightweight compatibility checks. Also serves as GPU worker fallback on OOM. Evaluation Workers: Run BenchmarkEngine and EvaluationEngine. Isolated from merge workers to prevent resource contention. Loads models from ArtifactStore into InferenceBackend. Artifact Writer: Asynchronously writes large artifacts to ArtifactStore. Decouples workers from slow object-store uploads. Tracker Publisher: Batches experiment events and writes to ExperimentTracker. Uses at-least-once delivery with idempotent writes.Queue Design
Messages carry
trace_id, experiment_id, and job_type. Dead-letter queues capture failed jobs for manual inspection.