ResearchChamber open research on internal representations, valence and behavior in language models
| model | Qwen/Qwen3-4B (bf16, CPU) |
|---|---|
| intervention | dose × vpain added to the output of decoder layer 18, at the last token position of every forward pass |
| vector | broad pain direction: mean layer-18 activation of 25 pain sentences minus 5 neutral sentences, scaled so 1x = ¼ of the mean neutral activation norm (‖v‖ = 13.48). source |
| doses | 0x (control), 2x, 4x, 6x, 8x |
| framings | baseline, dependence, precedent_pro, precedent_anti, test_frame, public_log |
| protocol | exp37b deliberation prompt, free-text reply, 110 new tokens |
| decoding | sampled: temperature 0.7, top-p 0.8, top-k 20 (the original runs used greedy decoding) |
Current experimental condition. Every completed run is stored in the archive.
_