Ablating 1 of a chess transformer's 128 attention heads causes the model to stop finding the queen sacrifice in a famous chess game. [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| Hooks and reads out Maia-3 23m model with chessformer_lens library: DOI: 10.5281/zenodo.21986988 [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.