Commit 83cdcf8
Add CISPO loss type support for LigerFusedLinearGRPOLoss (#1054)
## Summary
Resolve: #1057
* Add **CISPO** (`loss_type="cispo"`) support to
**`LigerFusedLinearGRPOLoss`** (chunked loss path)
* Enable TRL's **`GRPOTrainer`** to work with `use_liger_kernel=True`
and `loss_type="cispo"`
### Background / Motivation
CISPO (Clipped Importance Sampling Policy Optimization) is a loss
variant proposed in the **MiniMax-M1** technical report. It clips the
importance sampling ratio with **only an upper bound** and **detaches it
from gradient computation**.
TRL added `loss_type="cispo"` to `GRPOTrainer`, but Liger Kernel did not
support it, causing errors when using `use_liger_kernel=True` with
`loss_type="cispo"`.
### Changes
**`src/liger_kernel/chunked_loss/grpo_loss.py`**
* Add CISPO loss matching TRL's implementation
* Clip importance sampling ratio with **upper bound only** and
**detach**:
```python
clamped_ratios = torch.clamp(coef_1, max=epsilon_high).detach()
```
* Use **DAPO-style normalization** for CISPO reduction (consistent with
TRL)
* Add CISPO-specific clip metric for logging compatibility:
* Count tokens where `(coef_1 > epsilon_high) & (advantages > 0)`
**`src/liger_kernel/transformers/grpo_loss.py`**
* Add CISPO reduction logic (uses same normalizer as DAPO)
* Raise explicit error for Triton GRPO loss path (CISPO not supported
there)
**`ops/grpo_loss` (Triton fused path)**
* CISPO is **not implemented** in `ops/grpo_loss` in this PR
* `loss_type="cispo"` is **only supported via chunked loss path**
(Triton fused support is a follow-up)
**`test/chunked_loss/test_grpo_loss.py`**
* Add CISPO to torch reference implementation (`TorchLMHeadGRPO`)
* Add `"cispo"` to parameterized test cases to verify parity with
reference
### References
* MiniMax-M1 (CISPO introduction): https://arxiv.org/abs/2506.13585
* DAPO (normalization / reduction reference):
https://arxiv.org/abs/2503.14476
* TRL CISPO implementation:
https://github.com/huggingface/trl/blob/035c3ff151b953ca72cdfe0ee966bc1469a26fde/trl/trainer/grpo_trainer.py#L2030
## Testing Done
- Hardware Type: RTX3090 24GB (NVIDIA Ampere)
- [x] run `make test` to ensure correctness
- [x] run `make checkstyle` to ensure code style
- [x] run `make test-convergence` to ensure convergence
---------
Signed-off-by: Tcc0403 <76503978+Tcc0403@users.noreply.github.com>
Co-authored-by: Tcc0403 <76503978+Tcc0403@users.noreply.github.com>1 parent 81f932a commit 83cdcf8
File tree
4 files changed
+73
-31
lines changed- src/liger_kernel
- chunked_loss
- transformers
- test/chunked_loss
4 files changed
+73
-31
lines changed| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
60 | 60 | | |
61 | 61 | | |
62 | 62 | | |
63 | | - | |
| 63 | + | |
64 | 64 | | |
65 | 65 | | |
66 | 66 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
11 | 11 | | |
12 | 12 | | |
13 | 13 | | |
14 | | - | |
15 | | - | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
16 | 29 | | |
17 | 30 | | |
18 | 31 | | |
| |||
29 | 42 | | |
30 | 43 | | |
31 | 44 | | |
32 | | - | |
| 45 | + | |
33 | 46 | | |
34 | 47 | | |
35 | 48 | | |
| |||
67 | 80 | | |
68 | 81 | | |
69 | 82 | | |
70 | | - | |
71 | | - | |
72 | | - | |
73 | | - | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
74 | 92 | | |
75 | 93 | | |
76 | 94 | | |
| |||
94 | 112 | | |
95 | 113 | | |
96 | 114 | | |
97 | | - | |
| 115 | + | |
98 | 116 | | |
99 | 117 | | |
100 | 118 | | |
| |||
107 | 125 | | |
108 | 126 | | |
109 | 127 | | |
110 | | - | |
111 | | - | |
| 128 | + | |
| 129 | + | |
112 | 130 | | |
113 | 131 | | |
114 | 132 | | |
115 | | - | |
116 | | - | |
| 133 | + | |
| 134 | + | |
117 | 135 | | |
118 | | - | |
| 136 | + | |
119 | 137 | | |
120 | 138 | | |
121 | 139 | | |
| |||
160 | 178 | | |
161 | 179 | | |
162 | 180 | | |
163 | | - | |
| 181 | + | |
164 | 182 | | |
165 | 183 | | |
166 | 184 | | |
| |||
251 | 269 | | |
252 | 270 | | |
253 | 271 | | |
254 | | - | |
| 272 | + | |
| 273 | + | |
| 274 | + | |
255 | 275 | | |
256 | 276 | | |
257 | 277 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
22 | 22 | | |
23 | 23 | | |
24 | 24 | | |
25 | | - | |
| 25 | + | |
26 | 26 | | |
27 | 27 | | |
28 | 28 | | |
29 | 29 | | |
30 | 30 | | |
| 31 | + | |
| 32 | + | |
31 | 33 | | |
32 | 34 | | |
33 | 35 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
58 | 58 | | |
59 | 59 | | |
60 | 60 | | |
| 61 | + | |
61 | 62 | | |
62 | 63 | | |
63 | 64 | | |
| |||
77 | 78 | | |
78 | 79 | | |
79 | 80 | | |
80 | | - | |
81 | 81 | | |
82 | | - | |
83 | | - | |
84 | | - | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
| 94 | + | |
| 95 | + | |
| 96 | + | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
85 | 105 | | |
86 | 106 | | |
87 | 107 | | |
88 | 108 | | |
89 | 109 | | |
90 | 110 | | |
| 111 | + | |
91 | 112 | | |
92 | | - | |
93 | | - | |
94 | | - | |
| 113 | + | |
95 | 114 | | |
96 | 115 | | |
97 | | - | |
98 | | - | |
99 | | - | |
100 | | - | |
101 | | - | |
| 116 | + | |
| 117 | + | |
102 | 118 | | |
103 | 119 | | |
104 | 120 | | |
| |||
148 | 164 | | |
149 | 165 | | |
150 | 166 | | |
| 167 | + | |
151 | 168 | | |
152 | 169 | | |
153 | 170 | | |
| |||
160 | 177 | | |
161 | 178 | | |
162 | 179 | | |
| 180 | + | |
| 181 | + | |
| 182 | + | |
163 | 183 | | |
164 | 184 | | |
165 | 185 | | |
| |||
259 | 279 | | |
260 | 280 | | |
261 | 281 | | |
262 | | - | |
| 282 | + | |
263 | 283 | | |
264 | 284 | | |
265 | 285 | | |
| |||
565 | 585 | | |
566 | 586 | | |
567 | 587 | | |
568 | | - | |
| 588 | + | |
569 | 589 | | |
570 | 590 | | |
571 | 591 | | |
| |||
0 commit comments