The Salt - Curated AI
Subscribe
Sign in
Home
Notes
AI Notebooks
AI Repositories
Related Articles
deep dive
Archive
About
Latest
Top
Discussions
Lossless Diffusion Speedups, Random Attention, and One-Shot Distillation
The Weekly Salt #127
Sep 10
•
Benjamin Marie
2
August 2026
Hyperparameter Transfer, Agent Optimization, and Efficient MoE Deployment
The Weekly Salt #126
Aug 27
•
Benjamin Marie
1
Restoring KV Caches, Reducing Overthinking, and Diagnosing ALiBi
The Weekly Salt #125
Aug 7
•
Benjamin Marie
2
July 2026
Rich Feedback Outperforms Scalar Rewards on Open-Ended Tasks
The Weekly Salt #124
Jul 23
•
Benjamin Marie
4
Gemma 4: A Technical Look at the Architecture and Training
An efficient global-local attention architecture
Jul 16
•
Benjamin Marie
5
Verifying, Morphing, and Reader-Testing LLMs
The Weekly Salt #123
Jul 8
•
Benjamin Marie
2
June 2026
LoRA's Scaling Factor (Alpha): Still Misunderstood?
The Weekly Salt #122
Jun 18
•
Benjamin Marie
2
Flow-Based Token Credit for Reasoning RL
The Weekly Salt #120
Jun 11
•
Benjamin Marie
2
May 2026
Collaborative Parallel Thinking for Efficient Test-Time Scaling
The Weekly Salt #119
May 28
•
Benjamin Marie
Agents Fail to Reject Stale Memories
The Weekly Salt #118
May 20
•
Benjamin Marie
2
DeepSeek-V4: The Interesting Part Is the Attention Architecture
CSA, HCA, shared KV, mHC, ... How to make a good and efficient model with 1 million tokens in context
May 12
•
Benjamin Marie
1
April 2026
Expert Masking for More Efficient Expert Offloading
The Weekly Salt #117
Apr 29
•
Benjamin Marie
1
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts