- Training isn’t done at 4-bits, to date this small size has only been for infer...

acchow · 2024-03-19T00:48:56 1710809336

Intuitively, there is a ton of redundancy and we still have a long way we can still compress things.

imtringued · 2024-03-19T08:07:30 1710835650

Each token is represented by a vector of 4096 floats. Of course there is redundancy.

tmalsburg2 · 2024-03-19T08:53:32 1710838412

> - Training isn’t done at 4-bits, to date this small size has only been for inference.

Wasn't there a paper from Microsoft two weeks ago or so where they trained on log₂(3) bits?

Edit: https://arxiv.org/pdf/2402.17764.pdf

terramex · 2024-03-19T11:16:01 1710846961

They don't "train on log₂(3) bit". Gradients and activations are still calculated at full (8-bit) precision and weights are quantised after every update.

This makes network minimise loss not only with regard to expected outcome but also minimises loss resulting from quantisation. With big networks their "knowledge" is encoded in relationships between weights, not in their absolute values so lower precision work well as long as network is big enough.

coffeebeqn · 2024-03-19T01:49:23 1710812963

Maybe the rounding errors are noise that is somewhat useful in a big enough neutral net. Image generators also generate noise to work on