Notation
Weight Tensor Shapes for LLMs
For a transformer with hidden size H, vocab size V, intermediate size I, attention heads A, and KV heads K:Parameter Deltas and Task Vectors
A parameter delta is the difference between two parameter sets:Norms
The L2 norm of a flattened parameter tensor:Precision Effects
All merge operations should accumulate in fp32 and cast to the target dtype at the end. This is an engineering assumption: fp16 accumulation causes measurable drift in merged models.
Cross-Links
- SLERP Math for spherical interpolation formulas.
- TIES Math for trim, elect, and disjoint merge.
- DARE Math for sparsification and rescaling.
- Task Arithmetic for task vector combinations.
- Tensor Math for broadcasting and elementwise rules.