Packed trits vs the bf16 master weights, and bf16 vs f32 scales (data inside)

#48
by dsbusiness20206 - opened

We compared this packed checkpoint with microsoft/bitnet-b1.58-2B-4T-bf16 and microsoft/bitnet-b1.58-2B-4T-gguf at their current revisions, using independent decoders. Two observations:

  1. Same trits as the GGUF, same scale at lower precision. In layer 0 q_proj and down_proj (24.2M weights) the trits here and in the I2_S GGUF are identical. In all 210 ternary tensors, weight_scale here is the GGUF's f32 scale rounded to bf16 (for layer 0 down_proj: 2.15625 vs 2.1631613), so the two differ by 0.14% at the median and 0.36% at most.
  2. These trits cannot be recomputed from the bf16 master weights. transformers WeightQuant (float32 mean) applied to the bf16 weights gives different trits for 1.22% of layer 0 q_proj and 0.57% of layer 0 down_proj weights. In each tensor all of them have the one bf16 value nearest the threshold, |w| = 0.5 × weight_scale, and this checkpoint maps that value sometimes to ±1 and sometimes to 0 (79,719 of 163,223 weights are ±1 in q_proj). That fits trits computed from higher-precision master weights, of which the bf16 repository is a rounded copy.

Is the f32 scale in the GGUF the reference, and should the bf16 repository (online mode) be expected to reproduce these trits?

Details, revisions and a standalone reproduction (Python standard library only): https://github.com/microsoft/BitNet/issues/632

Write-up and data: https://github.com/dmitrii-f-t27/trinity-memory/blob/master/docs/ternary-check.md

Dmitrii Fedorov, Trinity (GitHub: dmitrii-f-t27)

Sign up or log in to comment