Glossary › deepseek-r
GLOSSARY
deepseek-r
appears in 3 paper titles
Definition
Refers to DeepSeek-R1, the reasoning-focused model released by the Chinese lab DeepSeek; titles sometimes split the name and leave a bare "deepseek-r". Its distinguishing feature is that long chain-of-thought behaviour was elicited mainly through reinforcement learning against verifiable answers, rather than by imitating human-written reasoning traces. Because the weights were published openly, it became the standard reference point for reproducing RL-trained reasoning.
Explainers using this term
- Designing Distillation Data — Deciding What to Ask the TeacherDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning