An important caveat to that explanation is that the mixture of experts gets re-evaluated for every single token in the input sequence and not for high level tasks as in the example. There will be certain “experts” for looking at indentation tokens, or ones that look at specific word beginnings/prefixes etc.
An important caveat to that explanation is that the mixture of experts gets re-evaluated for every single token in the input sequence and not for high level tasks as in the example. There will be certain “experts” for looking at indentation tokens, or ones that look at specific word beginnings/prefixes etc.