Post
32
I think it's clear in retrospect that "frankenmerges", which repeated blocks of layers, amounted to a crude approximation of looped transformers architecture, hence them able to work at all instead of just breaking. They lucked out due to much of the signal passing through residual streams being preserved and only modulated along the way. That said, not all models are suited for this. Models which feature ever-increasing magnitudes as inference progressess through layers risk exploding precision limits.