r/MachineLearning · · 1 min read

Ablating 1 of a chess transformer's 128 attention heads causes the model to stop finding the queen sacrifice in a famous chess game. [P]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

Ablating 1 of a chess transformer's 128 attention heads causes the model to stop finding the queen sacrifice in a famous chess game. [P]

Hooks and reads out Maia-3 23m model with chessformer_lens library:
github.com/chessformer-lens/chessformer_lens

DOI: 10.5281/zenodo.21986988

submitted by /u/Weird-Asparagus4136
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning