A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Artificial Intelligence
Title:A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
Abstract:A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests as an evidence lattice and validate the geometric branch on a known affine S_5 carrier: all held-source folds pass one-step, composition, inverse, decoding, and commutativity gates. Structured curvature and held-domain conjugacy raise error monotonically, but only 23/30 strongest cells flip a closure gate, bounding rather than universalizing calibration. In post-trained Qwen/Qwen3-4B, frozen final-token h28 affine maps have mean held-entity error .519, versus .398 for within-test-domain cross-fit. Seven randomized entity splits and map geometry do not support a purely entity-specific account. Earlier h4/h16 layers fit one-step transitions better, but h4 conflict-state decoding is weak and lexical controls remain unresolved. Three matched intervention datasets regenerated from one frozen checkpoint show causal effects only at h28/h36. Outcome-aware refitting improves h28 one-step error to .474 (.469 with weighting), yet no refit passes composition. Learned finite worlds likewise preserve relative algebraic signals or shared charts without held-source affine closure. Within the tested carriers, state availability, causal use, local geometry, and reusable closure are separable. The result is limited to one pretrained model, sampled final-token layers, two finite worlds, and the tested affine or diagnostic function classes.
| Comments: | 12 pages, 7 figures, 4 tables; includes supplementary results and ancillary reproducibility files |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2608.13626 [cs.AI] |
| (or arXiv:2608.13626v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2608.13626
arXiv-issued DOI via DataCite (pending registration)
|
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
Ancillary files (details):
- ARXIV_SNAPSHOT_README.txt
- FILE_MANIFEST.sha256
- README.md
- SOURCE_COMMIT.txt
- configs/phase0/qwen3_4b_nnsight_smoke.yaml
- configs/phase1/qwen3_4b_alchemy_paired_intervention.yaml
- configs/phase1/qwen3_4b_alchemy_paired_intervention_h16.yaml
- configs/phase1/qwen3_4b_alchemy_paired_intervention_h28.yaml
- configs/phase1/qwen3_4b_alchemy_paired_intervention_h28_seed2.yaml
- configs/phase1/qwen3_4b_alchemy_paired_intervention_h28_seed3.yaml
- configs/phase1/qwen3_4b_alchemy_paired_intervention_h36.yaml
- configs/phase1/qwen3_4b_alchemy_probe.yaml
- configs/phase1/qwen3_4b_alchemy_probe_direct.yaml
- configs/phase1/qwen3_4b_alchemy_probe_focused.yaml
- configs/phase2/qwen3_4b_grounded_operator_causal.yaml
- configs/phase2/qwen3_4b_grounded_operator_causal_seed2.yaml
- configs/phase2/qwen3_4b_grounded_operator_causal_seed3.yaml
- configs/phase2/qwen3_4b_grounded_operator_fit.yaml
- configs/phase2/qwen3_4b_grounded_operator_fit_alpha_expanded.yaml
- configs/phase2/qwen3_4b_grounded_operator_fit_seed2.yaml
- configs/phase2/qwen3_4b_grounded_operator_fit_seed3.yaml
- configs/phase2/qwen3_4b_grounded_operator_mlp_upper.yaml
- configs/phase3/qwen3_4b_h2_composition_seed1.yaml
- configs/phase3/qwen3_4b_h2_composition_seed2.yaml
- configs/phase3/qwen3_4b_h2_composition_seed3.yaml
- configs/phase3/qwen3_4b_h3_commutativity_seed1.yaml
- configs/phase3/qwen3_4b_h3_commutativity_seed2.yaml
- configs/phase3/qwen3_4b_h3_commutativity_seed3.yaml
- configs/phase3/qwen3_4b_phase3_5_construct_smoke.yaml
- configs/phase3/qwen3_4b_phase3_5_construct_validation.yaml
- configs/phase3/qwen3_4b_phase3_5b_manipulation_calibration.yaml
- configs/phase4/synthetic_trajectories.yaml
- configs/phase4/synthetic_trajectories_smoke.yaml
- configs/phase4b/stage_ordering.yaml
- configs/phase4b/stage_ordering_smoke.yaml
- configs/phase4c/s5_cross_world.yaml
- configs/phase4c/s5_cross_world_smoke.yaml
- configs/phase4d/s5_operator_ladder.yaml
- configs/phase4d/s5_operator_ladder_reproduction_shard.yaml
- configs/phase4d/s5_operator_ladder_smoke.yaml
- configs/phase4e/s5_operator_scaling.yaml
- configs/phase4e/s5_operator_scaling_reproduction_shard.yaml
- configs/phase4e/s5_operator_scaling_smoke.yaml
- configs/phase4f/s5_aspect_ratio.yaml
- configs/phase4f/s5_aspect_ratio_reproduction_shard.yaml
- configs/phase4f/s5_aspect_ratio_smoke.yaml
- configs/phase6/reviewer_calibration.yaml
- configs/phase6/reviewer_calibration_smoke.yaml
- configs/phase7/layer_local_composition.yaml
- configs/phase8/reviewer_v2_diagnostics.yaml
- configs/phase8/reviewer_v2_diagnostics_smoke.yaml
- configs/phase8b/metric_attribution.yaml
- paper/claim-intent-manifest.json
- paper/terminology-ledger.md
- pyproject.toml
- reports/2026-08-08-phase0-qwen3-4b.md
- reports/2026-08-08-phase1-qwen3-4b.md
- reports/2026-08-08-phase2-grounded-operators.md
- reports/2026-08-08-phase3-5-construct-validation.md
- reports/2026-08-08-phase3-5b-manipulation-calibration.md
- reports/2026-08-08-phase3-h2-composition.md
- reports/2026-08-08-phase3-h3-commutativity.md
- reports/2026-08-08-phase4-ground-truth-trajectories.md
- reports/2026-08-08-phase4-postrun-audit.json
- reports/2026-08-09-phase4b-postrun-audit.json
- reports/2026-08-09-phase4b-stage-ordering.md
- reports/2026-08-09-phase4c-postrun-audit.json
- reports/2026-08-09-phase4c-s5-cross-world.md
- reports/2026-08-09-phase4d-postrun-audit.json
- reports/2026-08-09-phase4d-s5-operator-ladder.md
- reports/2026-08-09-phase4e-s5-operator-scaling.md
- reports/2026-08-09-phase4f-s5-aspect-ratio.md
- reports/2026-08-11-phase6-reviewer-calibration-audit.json
- reports/2026-08-11-phase7-layer-local-composition-audit.json
- reports/2026-08-11-phase7-layer-local-composition.md
- reports/2026-08-12-phase8-reviewer-v2-local-audit.json
- reports/2026-08-12-phase8-reviewer-v2-remote-audit.json
- reports/2026-08-12-phase8b-metric-attribution-local-audit.json
- reports/2026-08-12-phase8b-metric-attribution-remote-audit.json
- reports/audits/phase4e-s5-20260810T013925Z-independent-audit.json
- reports/audits/phase4f-s5-20260810T030514Z-independent-audit.json
- reports/phase3-5-construct-validation-preregistration.md
- reports/phase3-5b-manipulation-calibration-preregistration.md
- reports/phase3-h2-preregistration.md
- reports/phase3-h3-preregistration.md
- reports/phase4-ground-truth-trajectories-preregistration.md
- reports/phase4b-stage-ordering-preregistration.md
- reports/phase4c-s5-cross-world-preregistration.md
- reports/phase4d-s5-operator-ladder-preregistration.md
- reports/phase4e-s5-operator-scaling-preregistration.md
- reports/phase4f-s5-aspect-ratio-preregistration.md
- reports/phase6-reviewer-calibration-preregistration.md
- reports/phase7-layer-local-composition-preregistration.md
- reports/phase8-reviewer-v2-diagnostics-preregistration.md
- reports/phase8b-metric-attribution-analysis-plan.md
- requirements/dev.txt
- requirements/remote.txt
- scripts/analyze_phase0_reproducibility.py
- scripts/analyze_phase3_5.py
- scripts/analyze_phase3_h2.py
- scripts/analyze_phase3_h3.py
- scripts/analyze_phase4.py
- scripts/audit_phase4b_stage_ordering.py
- scripts/audit_phase4c_s5_cross_world.py
- scripts/audit_phase4d_s5_operator_ladder.py
- scripts/audit_phase4e_s5_operator_scaling.py
- scripts/audit_phase4f_s5_aspect_ratio.py
- scripts/audit_phase5_latex_pdf.py
- scripts/audit_phase5_manuscript.py
- scripts/audit_phase6_reviewer_calibration.py
- scripts/audit_phase7_layer_local_composition.py
- scripts/audit_phase8_reviewer_v2_diagnostics.py
- scripts/audit_phase8b_metric_attribution.py
- scripts/build_phase5_pdf.sh
- scripts/finalize_artifact_manifest.py
- scripts/generate_phase5_latex.py
- scripts/merge_phase4d_s5_operator_shards.py
- scripts/merge_phase4e_s5_operator_shards.py
- scripts/merge_phase4f_s5_aspect_ratio_shards.py
- scripts/phase0_nnsight_smoke.py
- scripts/phase0_state_probes_legacy_smoke.py
- scripts/phase1_paired_intervention.py
- scripts/phase1_probe_scan.py
- scripts/phase2_fit_operators.py
- scripts/phase2_mlp_upper_bound.py
- scripts/phase2_operator_causal.py
- scripts/phase3_5_construct_validation.py
- scripts/phase3_5b_manipulation_calibration.py
- scripts/phase3_h2_composition.py
- scripts/phase3_h3_commutativity.py
- scripts/phase4_synthetic_trajectories.py
- scripts/phase4b_stage_ordering.py
- scripts/phase4c_s5_cross_world.py
- scripts/phase4d_s5_operator_ladder.py
- scripts/phase4e_s5_operator_scaling.py
- scripts/phase4f_s5_aspect_ratio.py
- scripts/phase6_reviewer_calibration.py
- scripts/phase7_layer_local_composition.py
- scripts/phase8_reviewer_v2_diagnostics.py
- scripts/phase8b_metric_attribution.py
- scripts/plot_phase5_h5_stitching.py
- scripts/plot_phase5_measurement_ladder.py
- scripts/plot_phase8_reviewer_v2.py
- scripts/plot_reviewer_calibration.py
- scripts/prepare_arxiv_submission.py
- source_data/README.txt
- source_data/h1_threshold_sensitivity.csv
- source_data/h2_threshold_sensitivity.csv
- source_data/h4_order_dz_sensitivity.csv
- source_data/h5_stitching.csv
- source_data/h5_stitching_metadata.json
- source_data/reviewer_calibration_composition.csv
- source_data/reviewer_calibration_depth.csv
- source_data/reviewer_calibration_grounded.csv
- source_data/reviewer_calibration_metadata.json
- source_data/reviewer_calibration_positive.csv
- source_data/reviewer_v2_entity_geometry.csv
- source_data/reviewer_v2_layer_geometry.csv
- source_data/reviewer_v2_lexical_controls.csv
- source_data/reviewer_v2_metadata.json
- source_data/reviewer_v2_metric_attribution.csv
- source_data/reviewer_v2_structured_calibration.csv
- src/verb_operator_algebra/__init__.py
- src/verb_operator_algebra/alchemy.py
- src/verb_operator_algebra/commutativity.py
- src/verb_operator_algebra/commutativity_data.py
- src/verb_operator_algebra/composition.py
- src/verb_operator_algebra/composition_data.py
- src/verb_operator_algebra/config.py
- src/verb_operator_algebra/construct_causal.py
- src/verb_operator_algebra/construct_data.py
- src/verb_operator_algebra/construct_validation.py
- src/verb_operator_algebra/layer_local_composition.py
- src/verb_operator_algebra/manipulation_calibration.py
- src/verb_operator_algebra/metric_attribution.py
- src/verb_operator_algebra/operator_causal.py
- src/verb_operator_algebra/operator_data.py
- src/verb_operator_algebra/operator_fit.py
- src/verb_operator_algebra/phase0.py
- src/verb_operator_algebra/phase1_intervention.py
- src/verb_operator_algebra/phase1_probe.py
- src/verb_operator_algebra/reporting.py
- src/verb_operator_algebra/reviewer_calibration.py
- src/verb_operator_algebra/reviewer_v2_diagnostics.py
- src/verb_operator_algebra/s5_aspect_ratio.py
- src/verb_operator_algebra/s5_operator_diagnostics.py
- src/verb_operator_algebra/s5_operator_scaling.py
- src/verb_operator_algebra/s5_stage_analysis.py
- src/verb_operator_algebra/s5_world.py
- src/verb_operator_algebra/synthetic_analysis.py
- src/verb_operator_algebra/synthetic_model.py
- src/verb_operator_algebra/synthetic_stage_analysis.py
- src/verb_operator_algebra/synthetic_world.py
- tests/test_alchemy.py
- tests/test_audit_phase4b_stage_ordering.py
- tests/test_audit_phase4c_s5_cross_world.py
- tests/test_audit_phase4d_s5_operator_ladder.py
- tests/test_audit_phase4e_s5_operator_scaling.py
- tests/test_audit_phase4f_s5_aspect_ratio.py
- tests/test_audit_phase6_reviewer_calibration.py
- tests/test_audit_phase7_layer_local_composition.py
- tests/test_audit_phase8_reviewer_v2_diagnostics.py
- tests/test_audit_phase8b_metric_attribution.py
- tests/test_commutativity.py
- tests/test_commutativity_data.py
- tests/test_composition.py
- tests/test_composition_data.py
- tests/test_config.py
- tests/test_construct_data.py
- tests/test_construct_validation.py
- tests/test_finalize_artifact_manifest.py
- tests/test_layer_local_composition.py
- tests/test_manipulation_calibration.py
- tests/test_merge_phase4d_s5_operator_shards.py
- tests/test_merge_phase4e_s5_operator_shards.py
- tests/test_merge_phase4f_s5_aspect_ratio_shards.py
- tests/test_metric_attribution.py
- tests/test_operator_causal.py
- tests/test_operator_data.py
- tests/test_operator_fit.py
- tests/test_phase1_intervention.py
- tests/test_phase4b_config.py
- tests/test_phase4c_config.py
- tests/test_phase4d_config.py
- tests/test_phase4e_config.py
- tests/test_phase4f_config.py
- tests/test_phase4f_s5_aspect_ratio.py
- tests/test_phase6_config.py
- tests/test_prepare_arxiv_submission.py
- tests/test_reporting.py
- tests/test_reviewer_calibration.py
- tests/test_reviewer_v2_diagnostics.py
- tests/test_s5_operator_diagnostics.py
- tests/test_s5_operator_scaling.py
- tests/test_s5_stage_analysis.py
- tests/test_s5_world.py
- tests/test_synthetic_analysis.py
- tests/test_synthetic_stage_analysis.py
- tests/test_synthetic_world.py
- verb-operator-algebra-research-plan.md
Current browse context:
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
Recipes for Steering and Scaling LLMs via Sampling
Aug 28
-
Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention
Aug 28
-
FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
Aug 28
-
Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores
Aug 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.