Timeline
Post
Remote status
Replies
7
@lain is that even related? They already know which "subnetwork" is "tuned to similar performance", by way of expert activation. You can isolate which experts are activated for coding, no?
They just took those and put them in a smaller model to reduce overhead...
They just took those and put them in a smaller model to reduce overhead...
@WandererUber every expert in a3b is 3b. so a 2.6b model can't just take the experts. they must have derived more fundamental experts / pruned the useless nodes to achieve this.
Think it's legit?
@HAPPYPOTAMUS "I should finally make my local coding performance benchmark suite"
@WandererUber ah, i misread it! very interesting idea!