Artificial Analysis [2026b]
Artificial Analysis.
Kimi k2.7 code: Intelligence, performance & price analysis.
https://artificialanalysis.ai/models/kimi-k2-7-code,
2026b.
Open-weights (Moonshot AI), released June 12, 2026; Artificial
Analysis Intelligence Index score of 42; $0.95/$4.00 per 1M input/output
tokens.
Artificial Analysis [2026c]
Artificial Analysis.
Claude opus 4.8 (max): Intelligence, performance & price analysis.
https://artificialanalysis.ai/models/claude-opus-4-8,
2026c.
Claude Opus 4.8 (Adaptive Reasoning, Max Effort), released May 28,
2026; Artificial Analysis Intelligence Index score of 56.
Baek et al. [2023]
Jinheon Baek, Alham Fikri Aji, and Amir Saffari.
Knowledge-augmented language model prompting for zero-shot knowledge
graph question answering.
In Proceedings of the 1st Workshop on Natural Language
Reasoning and Structured Explanations (NLRSE), 2023.
KAPING: retrieves and verbalizes relevant knowledge-graph triples
into the LLM prompt for zero-shot knowledge-graph question answering.
Bauer et al. [2015]
Peter Bauer, Alan Thorpe, and Gilbert Brunet.
The quiet revolution of numerical weather prediction.
Nature, 525(7567):47–55, 2015.
Documents roughly one forecast-day-of-skill-per-decade gains and the
compute/assimilation basis of NWP.
Bengio et al. [2025a]
Yoshua Bengio, Michael K. Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro
Greiner, Matt MacDermott, Sören Mindermann, Adam Oberman, Jesse
Richardson, Oliver Richardson, Marc-Antoine Rondeau, Pierre-Luc St-Charles,
and David Williams-King.
Superintelligent agents pose catastrophic risks: Can Scientist AI
offer a safer path?, 2025a.
Proposes a non-agentic “Scientist AI”: a world model that generates
theories to explain data plus a question-answering inference machine, both
carrying explicit uncertainty, usable as a run-time safety guardrail rather
than an actor.
Bengio et al. [2025b]
Yoshua Bengio, Michael K. Cohen, Nikolay Malkin, Matt MacDermott, Damiano
Fornasiere, Pietro Greiner, and Younesse Kaddar.
Can a Bayesian oracle prevent harm from an agent?
In Proceedings of Machine Learning Research, volume 286,
2025b.
Derives run-time, context-dependent bounds on the probability an
action violates a safety specification, using Bayesian posteriors over
world-model hypotheses; arXiv:2408.05284.
Bernardo and Smith [2000]
José M. Bernardo and Adrian F. M. Smith.
Bayesian Theory.
Wiley, 2000.
M-closed / M-complete / M-open taxonomy of inference settings.
Bi et al. [2023]
Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian.
Accurate medium-range global weather forecasting with 3D neural
networks.
Nature, 619(7970):533–538, 2023.
Pangu-Weather: 3D neural weather model trained on reanalysis.
Chiatti et al. [2026]
Agnese Chiatti, Michael Cochez, Cristina Cornelio, Sebastijan
Dumančić, Artur d’Avila Garcez, Luis C. Lamb, Lia Morra, Mathias
Niepert, Robert Peharz, Alberto Speranzon, Maarten Stol, Annette ten Teije,
Thiviyan Thanapalasingam, Frank van Harmelen, Emile van Krieken, Antonio
Vergari, and Benjie Wang.
The RAIL principles for neurosymbolic AI: Reasoning, assurances,
interfacing and learning.
Communications of the ACM, 2026.
Result of Dagstuhl Seminar 25452; analyzes AI systems—including
physics-aware ML, DeepMind’s Alpha-* suite, causal learning, and
tool-augmented LLMs—along four neurosymbolic design axes (Reasoning,
Assurances, Interfacing, Learning), arguing that systematic neural/symbolic
integration, not scale alone, is the route to reliable AI in critical
domains.
Deutsch [2011]
David Deutsch.
The Beginning of Infinity: Explanations That Transform the
World.
Viking, 2011.
Source of the “hard-to-vary explanations” criterion.
Du et al. [2025]
Yufeng Du, Minyang Tian, Srikanth Ronanki, Subendhu Rongali, Sravan Bodapati,
Aram Galstyan, Azton Wells, Roy Schwartz, Eliu A. Huerta, and Hao Peng.
Context length alone hurts LLM performance despite perfect
retrieval.
In Findings of the Association for Computational Linguistics:
EMNLP 2025, 2025.
Shows input length itself, independent of retrieval quality,
substantially degrades LLM task performance (13.9%–85%) even when all
relevant evidence is perfectly retrievable and irrelevant tokens are masked.
Edge et al. [2024]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody,
Steven Truitt, and Jonathan Larson.
From local to global: A graph RAG approach to query-focused
summarization, 2024.
Farmer [2024]
J. Doyne Farmer.
Making Sense of Chaos: A Better Economics for a Better World.
Yale University Press, 2024.
Complexity-economics case for mechanistic, agent-based simulation of
the economy, as meteorology did; conjectures economic systems may be more
tractable than the weather.
FelloAI [2026]
FelloAI.
Qwen3.7-max review 2026: Benchmarks, pricing, verdict.
https://felloai.com/qwen-3-7-max-review/, 2026.
Alibaba Qwen3.7-Max, released May 20, 2026; Artificial Analysis
Intelligence Index score of 56.6; $2.50/$7.50 per 1M input/output tokens
(DashScope).
Friston et al. [2017]
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, and
Giovanni Pezzulo.
Active inference: A process theory.
Neural Computation, 29(1):1–49, 2017.
Gao et al. [2023]
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie
Callan, and Graham Neubig.
PAL: Program-aided language models.
In Proceedings of the 40th International Conference on Machine
Learning, volume 202 of Proceedings of Machine Learning Research,
pages 10764–10799, 2023.
Has the LLM generate a program as the reasoning trace but offloads
execution to a Python interpreter, decoupling solving (computation) from
decomposition (language); the canonical code-execution-as-computation-tool
augmentation.
Garcez [2025]
Artur d’Avila Garcez.
Neurosymbolic AI: Towards sound reasoning and causal learning and
the road to AGI.
https://www.staff.city.ac.uk/~aag/papers/NeSyAIGarcez2025,
2025.
Argues chain-of-thought prompting and post-hoc RLHF cannot fix LLM
hallucination because errors compound in continuous, ungrounded computation;
proposes the neurosymbolic cycle (extract, reason, distill) as a route to
reliable, data-efficient reasoning and AGI, with agentic AI becoming
neurosymbolic once code execution is paired with symbolic control.
Garcez and Lamb [2023]
Artur d’Avila Garcez and Luís C. Lamb.
Neurosymbolic AI: The 3rd wave.
Artificial Intelligence Review, 56:12387–12406,
2023.
Gelman and Yao [2021]
Andrew Gelman and Yuling Yao.
Holes in Bayesian statistics.
Journal of Physics G: Nuclear and Particle Physics,
48(1):014002, 2021.
doi: 10.1088/1361-6471/abc3a5.
Catalogues structural tensions in Bayesian inference, including the
“Cantor’s corner” argument (hole 6): checking a model against
data—necessary in practice whenever the model class is not known a priori
to contain the truth—is not itself a coherent Bayesian operation, since it
requires stepping outside the assumed model to entertain alternatives not
covered by the prior.
Harnad [1990]
Stevan Harnad.
The symbol grounding problem.
Physica D: Nonlinear Phenomena, 42(1–3):335–346, 1990.
doi: 10.1016/0167-2789(90)90087-6.
Canonical statement of the symbol-grounding problem we argue does not
apply to our notion of grounding.
Hersbach et al. [2020]
Hans Hersbach et al.
The ERA5 global reanalysis.
Quarterly Journal of the Royal Meteorological Society,
146(730):1999–2049, 2020.
Single coherent best-estimate of the global atmospheric state.
Hoeting et al. [1999]
Jennifer A. Hoeting, David Madigan, Adrian E. Raftery, and Chris T. Volinsky.
Bayesian model averaging: A tutorial.
Statistical Science, 14(4):382–417, 1999.
Huh et al. [2024]
Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola.
The platonic representation hypothesis.
In International Conference on Machine Learning, 2024.
Hutter [2005]
Marcus Hutter.
Universal Artificial Intelligence: Sequential Decisions Based
on Algorithmic Probability.
Springer, 2005.
Defines AIXI, the unbounded Bayesian-optimal predictor/agent.
Kalnay [2003]
Eugenia Kalnay.
Atmospheric Modeling, Data Assimilation and Predictability.
Cambridge University Press, 2003.
Standard reference treating data assimilation (4D-Var, ensemble
Kalman filtering) as Bayesian conditioning of a physical model.
Karniadakis et al. [2021]
George Em Karniadakis, Ioannis G. Kevrekidis, Lu Lu, Paris Perdikaris, Sifan
Wang, and Liu Yang.
Physics-informed machine learning.
Nature Reviews Physics, 3(6):422–440,
2021.
Embeds known governing equations into neural-network training as soft
constraints; the canonical mechanism-regularized (but still neural,
unverified) alternative to an explicit program.
Kaufmann et al. [2025]
Rafael Kaufmann, Felix Neubürger, Michael Walters, Thomas Kopinski, and
Dimitrije Marković.
The CRISTAL method: Fast, reliable analytical problem-solving with
pre-synthesized grounded world models.
In Proceedings of the 19th Conference on Neurosymbolic Learning
and Reasoning (NeSy), Proceedings of Machine Learning Research, 2025.
Primordia / GAIA Lab. Introduces grounded world models (GWMs): a
synthesized, continually-refined probabilistic program enabling full Bayesian
inference; reaches Bayes-optimal accuracy on a synthetic-equities benchmark
with 5 examples where SOTA LLMs plateau near 40%.
Kaufmann et al. [2026]
Rafael Kaufmann, Harald Stromfelt, Thomas Minter, and Sandeep Ramesh.
Coherent world-model meshes: Cross-case joint inference for
structural causal prediction.
Companion paper (Paper 2). Develops the cross-case Case Mesh: a
shared-latent joint posterior across coupled cases, the cost of brute-forcing
cross-case coherence on a scale-free collision graph, and the decision-impact
benchmark for neglected downside correlation. Builds on and cites the present
paper., 2026.
Koller and Friedman [2009]
Daphne Koller and Nir Friedman.
Probabilistic graphical models: Principles and techniques.
MIT Press, 2009.
Koza [1992]
John R. Koza.
Genetic programming: On the programming of computers by means of
natural selection.
MIT Press, 1992.
Founding text of genetic programming: evolving computer programs
against a fitness measure rather than hand-writing them.
Kumar et al. [2025]
Akarsh Kumar, Jeff Clune, Joel Lehman, and Kenneth O. Stanley.
Questioning representational optimism in deep learning: The fractured
entangled representation hypothesis.
arXiv preprint arXiv:2505.11581, 2025.
Lam et al. [2023]
Remi Lam et al.
Learning skillful medium-range global weather forecasting.
Science, 382(6677):1416–1421, 2023.
GraphCast: ML weather emulator trained on ERA5 reanalysis.
LeCun [2022]
Yann LeCun.
A path towards autonomous machine intelligence.
Open Review preprint, 2022.
Lewis et al. [2020]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir
Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim
Rocktäschel, Sebastian Riedel, and Douwe Kiela.
Retrieval-augmented generation for knowledge-intensive NLP tasks.
Advances in Neural Information Processing Systems, 33, 2020.
Introduces retrieval-augmented generation (RAG): a non-parametric
retriever supplies passages that a generator conditions on, the canonical
long-context-retrieval augmentation of an LLM.
Liu et al. [2024]
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua,
Fabio Petroni, and Percy Liang.
Lost in the middle: How language models use long contexts.
Transactions of the Association for Computational Linguistics,
12:157–173, 2024.
doi: 10.1162/tacl_a_00638.
Shows LLM performance on long-context tasks degrades significantly,
and non-monotonically with position, as input context grows, even for models
explicitly built for long contexts.
Lorenz [1963]
Edward N. Lorenz.
Deterministic nonperiodic flow.
Journal of the Atmospheric Sciences, 20(2):130–141, 1963.
Sensitive dependence on initial conditions; origin of the finite
predictability horizon.
Lundberg and Lee [2017]
Scott M. Lundberg and Su-In Lee.
A unified approach to interpreting model predictions.
In Advances in Neural Information Processing Systems
(NeurIPS), volume 30, pages 4765–4774, 2017.
SHAP: Shapley-value attributions for individual predictions.
Meyerson et al. [2025]
Elliot Meyerson, Giuseppe Paolo, Roberto Dailey, Hormoz Shahrzad, Olivier
Francon, Conor F. Hayes, Xin Qiu, Babak Hodjat, and Risto Miikkulainen.
Solving a million-step LLM task with zero errors, 2025.
Quantifies the cost of per-step error correction in long-horizon
agentic execution: since per-step error rates do not vanish with scale,
expected cost to complete an -step task grows as absent
decomposition and voting-based verification, motivating extreme task
decomposition to make error correction and rework tractable.
Mlodozeniec et al. [2025]
Bruno Kacper Mlodozeniec, David Krueger, and Richard E. Turner.
Position: Probabilistic modelling is sufficient for causal inference.
In Proceedings of the 42nd International Conference on Machine
Learning (ICML), volume 267 of Proceedings of Machine Learning
Research, pages 81810–81840, 2025.
Argues any causal inference question can be answered within standard
probabilistic modelling and inference, reinterpreting causal-specific tools
(e.g. the do-operator) as emerging from probabilistic modelling on a suitably
expanded model rather than requiring bespoke causal notation.
Packer et al. [2023]
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion
Stoica, and Joseph E. Gonzalez.
MemGPT: Towards LLMs as operating systems.
arXiv preprint arXiv:2310.08560, 2023.
Virtual context management: a hierarchical, agent-managed memory
store that pages facts in and out of the LLM’s context across calls, the
canonical instance of an “agentic memory” store.
Pearl [2009]
Judea Pearl.
Causal inference in statistics: An overview.
Statistics Surveys, 3:96–146, 2009.
Quine [1951]
Willard Van Orman Quine.
Two dogmas of empiricism.
The Philosophical Review, 60(1):20–43,
1951.
Source of confirmation holism: statements face experience only as a
corporate body, not one by one, against which we read our posture that
“observables” are pragmatically defined by a model’s context of
applicability rather than by any privileged ontological status.
Ribeiro et al. [2016]
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin.
“why should I trust you?”: Explaining the predictions of any
classifier.
In Proceedings of the 22nd ACM SIGKDD International Conference
on Knowledge Discovery and Data Mining, pages 1135–1144, 2016.
LIME: local surrogate explanations of individual predictions.
Richens and Everitt [2024]
Jonathan Richens and Tom Everitt.
Robust agents learn causal world models.
In Proceedings of the 12th International Conference on Learning
Representations (ICLR), 2024.
Proves any agent satisfying a regret bound under a large set of
distributional (interventional) shifts must have learned an approximate
causal model of the data-generating process, converging to the true causal
model for optimal agents.
Ro et al. [2025]
Yeonju Ro, Haoran Qiu, Íñigo Goiri, Rodrigo Fonseca, Ricardo Bianchini,
Aditya Akella, Zhangyang Wang, Mattan Erez, and Esha Choukse.
Sherlock: Reliable and efficient agentic workflow execution.
2025.
Documents that errors in agentic workflows propagate and compound
across downstream steps, and quantifies the latency/cost overhead of the
verification and rollback (rework) required to catch them, motivating
selective, cost-optimal verification.
Schick et al. [2023]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria
Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom.
Toolformer: Language models can teach themselves to use tools.
Advances in Neural Information Processing Systems, 36, 2023.
Self-supervised training of an LLM to decide when to call external
tools (a calculator, Q&A system, search engine) and incorporate their
outputs into generation.
Schmidt and Lipson [2009]
Michael Schmidt and Hod Lipson.
Distilling free-form natural laws from experimental data.
Science, 324(5923):81–85, 2009.
doi: 10.1126/science.1165893.
Symbolic regression: searches jointly over equation form and
parameters to recover free-form analytical laws from experimental data.
Sobieski and Biecek [2024]
Bartosz Sobieski and Przemysław Biecek.
Global counterfactual directions.
In Proceedings of the European Conference on Computer Vision
(ECCV), 2024.
Discovers latent directions that flip a classifier’s decision across
an entire dataset—a global, model-level counterfactual extraction, in
contrast to local post-hoc attributions such as LIME and SHAP.
Solomonoff [1964]
Ray J. Solomonoff.
A formal theory of inductive inference, parts i and ii.
Information and Control, 7(1–2):1–22,
224–254, 1964.
Spirtes et al. [2000]
Peter Spirtes, Clark N. Glymour, and Richard Scheines.
Causation, Prediction, and Search.
MIT Press, 2nd edition, 2000.
Canonical reference for constraint-based causal discovery (the PC/FCI
algorithm family), which learns causal structure from observational data
rather than from expert-informed priors.
Vals AI [2026]
Vals AI.
Finance agent v2: Evaluating agents on core financial analyst tasks.
https://www.vals.ai/benchmarks/fabv2, 2026.
Benchmark of LLM agents on entry-level financial-analyst tasks over
public filings; no model clears and Claude Opus 4.8 scores
. The harness withholds code-execution and structured-memory
tools.
Walters et al. [2025a]
Michael Walters, Rafael Kaufmann, Justice Sefas, and Thomas Kopinski.
Free energy risk metrics for systemically safe AI: Gatekeeping
multi-agent study, 2025a.
Primordia / GAIA Lab. Introduces a Cumulative Risk Exposure metric
grounded in the Free Energy Principle for online, uncertainty-aware,
preference-based risk governance in agentic and multi-agent systems,
requiring only stakeholder-specified outcome preferences rather than
exhaustive world models.
Walters et al. [2025b]
Michael Walters, Thomas Kopinski, Rohil Rao, Rafael Kaufmann, and Alf
Köhn-Seemann.
The GAIA tech tree: A neurosymbolic AI framework for strategic
technological decision-making.
Working Paper, 2025b.
Wong et al. [2023]
Lionel Wong, Gabriel Grand, Alexander K. Lew, Noah D. Goodman, Vikash K.
Mansinghka, Jacob Andreas, and Joshua B. Tenenbaum.
From word models to world models: Translating from natural language
to the probabilistic language of thought, 2023.
LLM-driven translation of natural language into probabilistic
programs.
Yao et al. [2022]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan,
and Yuan Cao.
ReAct: Synergizing reasoning and acting in language models.
arXiv preprint arXiv:2210.03629, 2022.
Interleaves verbal reasoning traces with tool/environment actions
(e.g. search, code execution) in a single LLM agent loop.