The Golden Age of Open Models
The gap between proprietary models (like GPT-4) and open weights models is rapidly closing. Meta's Llama-3 and Mistral's models are leading the charge.
Llama-3 8B Performance
Llama-3 8B punches way above its weight class. It exhibits excellent reasoning, high-quality coding capabilities, and strict instruction following, though its context window is limited to 8K.
Mistral NeMo (12B)
Mistral NeMo, built in collaboration with Nvidia, offers a massive 128K context window. It excels in summarization and multilingual tasks, providing a strong alternative to Llama-3 for document-heavy workloads.