Spyke

Syndicated from the fediverse. Read and engage on the original instance.

View original on programming.dev

11 replies

it's by far best sub trillion parameter model. fits in 512gb with 1m context at q4. Or less rental hardware than any model that scores higher than it.

2

Are these any good? I mean, within their size class, obviously - not expecting them to compare to Kimi K3!

Models the size of that 135M one open up some interesting use cases for edge devices.

2
cyd
lemmy.world

Very likely the same architecture, just with more post training.

2

Just repeating rumors (sorry, should have been clearer): GLM 5.3 is 5.2 with extra post training, their big upcoming model is 5.5. It kind of makes sense too; pretraining on a new architecture/size is expensive and it's natural for these companies to try to wring an extra minor version or two out of each one.

4

You reached the end

GLM-5.3 Imminent drop | Spyke