The Salt - Curated AI
Subscribe
Sign in
Home
Notes
AI Notebooks
AI Repositories
Related Articles
deep dive
Archive
About
deep dive
Latest
Top
Discussions
Gemma 4: A Technical Look at the Architecture and Training
An efficient global-local attention architecture
Jul 16
•
Benjamin Marie
5
DeepSeek-V4: The Interesting Part Is the Attention Architecture
CSA, HCA, shared KV, mHC, ... How to make a good and efficient model with 1 million tokens in context
May 12
•
Benjamin Marie
1
A Review of GLM-5: 744B-Parameter LLM Built for 200K Context and Agentic Training
With Sparse and Multi-Head Latent Attention
Feb 26
•
Benjamin Marie
4
Qwen3-VL: DeepStack Fusion, Interleaved-MRoPE, and a Native 256K Interleaved Context Window
Understanding why Qwen3-VL are the best VLMs
Dec 18, 2025
•
Benjamin Marie
5
1
Jet-Nemotron: Searching for the Best Attention Architecture
DeltaNet + Hardware-aware Search
Sep 23, 2025
•
Benjamin Marie
3
1
Magistral: Advancing Reasoning with Efficient GRPO Training
No More KL Penalty, No Need for a Reference Model
Jun 12, 2025
•
Benjamin Marie
3
Qwen3 Technical Report: Reasoning in Pre-Training and Post-Training
Plus a Brief Look at the Limitations of the Multilingual Evaluation
May 16, 2025
•
Benjamin Marie
6
Qwen2.5-VL: High-Resolution Vision Encoding with Efficient Windowed Attention
Also impressive in language generation tasks!
Mar 6, 2025
•
Benjamin Marie
6
TÜLU 3: The Post-Training Recipe
SFT + DPO + RLVR
Dec 19, 2024
•
Benjamin Marie
5
TÜLU 3's High-Quality Synthetic Datasets for Post-Training LLMs
Made by GPT-4o
Dec 5, 2024
•
Benjamin Marie
4
Go Zero-Shot for Cheaper LLM Evaluations
Unless you use a generative benchmark
Nov 6, 2024
•
Benjamin Marie
4
1
Evaluating AdEMAMix: A New Optimizer for Faster, More Efficient LLM Training
But with hyperparameter values not easy to find!
Oct 9, 2024
•
Benjamin Marie
6
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts