Mixtral 8x22B: Large-Scale Sparse¶ Released: 2024 Physical: 141B parameters Active: ~39B per token Performance: Excellent on benchmarks Context: 32K tokens extended Use: High-performance inference