Primordia Co.Grounded World Models

References

  • Artificial Analysis [2026a] Artificial Analysis. GLM-5.2 is the new leading open weights model on the artificial analysis intelligence index. https://artificialanalysis.ai/articles/glm-5-2-is-the-new-leading-open-weights-model-on-the-artificial-analysis-intelligence-index, 2026a. Zhipu/Z.ai GLM-5.2 (744B total / 40B active, MIT license), released June 16, 2026; Artificial Analysis Intelligence Index score of 51; $1.40/$4.40 per 1M input/output tokens.
  • Artificial Analysis [2026b] Artificial Analysis. Kimi k2.7 code: Intelligence, performance & price analysis. https://artificialanalysis.ai/models/kimi-k2-7-code, 2026b. Open-weights (Moonshot AI), released June 12, 2026; Artificial Analysis Intelligence Index score of 42; $0.95/$4.00 per 1M input/output tokens.
  • Artificial Analysis [2026c] Artificial Analysis. Claude opus 4.8 (max): Intelligence, performance & price analysis. https://artificialanalysis.ai/models/claude-opus-4-8, 2026c. Claude Opus 4.8 (Adaptive Reasoning, Max Effort), released May 28, 2026; Artificial Analysis Intelligence Index score of 56.
  • Baek et al. [2023] Jinheon Baek, Alham Fikri Aji, and Amir Saffari. Knowledge-augmented language model prompting for zero-shot knowledge graph question answering. In Proceedings of the 1st Workshop on Natural Language Reasoning and Structured Explanations (NLRSE), 2023. KAPING: retrieves and verbalizes relevant knowledge-graph triples into the LLM prompt for zero-shot knowledge-graph question answering.
  • Bauer et al. [2015] Peter Bauer, Alan Thorpe, and Gilbert Brunet. The quiet revolution of numerical weather prediction. Nature, 525(7567):47–55, 2015. Documents roughly one forecast-day-of-skill-per-decade gains and the compute/assimilation basis of NWP.
  • Bengio et al. [2025a] Yoshua Bengio, Michael K. Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, Sören Mindermann, Adam Oberman, Jesse Richardson, Oliver Richardson, Marc-Antoine Rondeau, Pierre-Luc St-Charles, and David Williams-King. Superintelligent agents pose catastrophic risks: Can Scientist AI offer a safer path?, 2025a. Proposes a non-agentic “Scientist AI”: a world model that generates theories to explain data plus a question-answering inference machine, both carrying explicit uncertainty, usable as a run-time safety guardrail rather than an actor.
  • Bengio et al. [2025b] Yoshua Bengio, Michael K. Cohen, Nikolay Malkin, Matt MacDermott, Damiano Fornasiere, Pietro Greiner, and Younesse Kaddar. Can a Bayesian oracle prevent harm from an agent? In Proceedings of Machine Learning Research, volume 286, 2025b. Derives run-time, context-dependent bounds on the probability an action violates a safety specification, using Bayesian posteriors over world-model hypotheses; arXiv:2408.05284.
  • Bernardo and Smith [2000] José M. Bernardo and Adrian F. M. Smith. Bayesian Theory. Wiley, 2000. M-closed / M-complete / M-open taxonomy of inference settings.
  • Bi et al. [2023] Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3D neural networks. Nature, 619(7970):533–538, 2023. Pangu-Weather: 3D neural weather model trained on reanalysis.
  • Chiatti et al. [2026] Agnese Chiatti, Michael Cochez, Cristina Cornelio, Sebastijan Dumančić, Artur d’Avila Garcez, Luis C. Lamb, Lia Morra, Mathias Niepert, Robert Peharz, Alberto Speranzon, Maarten Stol, Annette ten Teije, Thiviyan Thanapalasingam, Frank van Harmelen, Emile van Krieken, Antonio Vergari, and Benjie Wang. The RAIL principles for neurosymbolic AI: Reasoning, assurances, interfacing and learning. Communications of the ACM, 2026. Result of Dagstuhl Seminar 25452; analyzes AI systems—including physics-aware ML, DeepMind’s Alpha-* suite, causal learning, and tool-augmented LLMs—along four neurosymbolic design axes (Reasoning, Assurances, Interfacing, Learning), arguing that systematic neural/symbolic integration, not scale alone, is the route to reliable AI in critical domains.
  • Deutsch [2011] David Deutsch. The Beginning of Infinity: Explanations That Transform the World. Viking, 2011. Source of the “hard-to-vary explanations” criterion.
  • Du et al. [2025] Yufeng Du, Minyang Tian, Srikanth Ronanki, Subendhu Rongali, Sravan Bodapati, Aram Galstyan, Azton Wells, Roy Schwartz, Eliu A. Huerta, and Hao Peng. Context length alone hurts LLM performance despite perfect retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2025, 2025. Shows input length itself, independent of retrieval quality, substantially degrades LLM task performance (13.9%–85%) even when all relevant evidence is perfectly retrievable and irrelevant tokens are masked.
  • Edge et al. [2024] Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. From local to global: A graph RAG approach to query-focused summarization, 2024.
  • Farmer [2024] J. Doyne Farmer. Making Sense of Chaos: A Better Economics for a Better World. Yale University Press, 2024. Complexity-economics case for mechanistic, agent-based simulation of the economy, as meteorology did; conjectures economic systems may be more tractable than the weather.
  • FelloAI [2026] FelloAI. Qwen3.7-max review 2026: Benchmarks, pricing, verdict. https://felloai.com/qwen-3-7-max-review/, 2026. Alibaba Qwen3.7-Max, released May 20, 2026; Artificial Analysis Intelligence Index score of 56.6; $2.50/$7.50 per 1M input/output tokens (DashScope).
  • Friston et al. [2017] Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, and Giovanni Pezzulo. Active inference: A process theory. Neural Computation, 29(1):1–49, 2017.
  • Gao et al. [2023] Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. PAL: Program-aided language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 10764–10799, 2023. Has the LLM generate a program as the reasoning trace but offloads execution to a Python interpreter, decoupling solving (computation) from decomposition (language); the canonical code-execution-as-computation-tool augmentation.
  • Garcez [2025] Artur d’Avila Garcez. Neurosymbolic AI: Towards sound reasoning and causal learning and the road to AGI. https://www.staff.city.ac.uk/~aag/papers/NeSyAIGarcez2025, 2025. Argues chain-of-thought prompting and post-hoc RLHF cannot fix LLM hallucination because errors compound in continuous, ungrounded computation; proposes the neurosymbolic cycle (extract, reason, distill) as a route to reliable, data-efficient reasoning and AGI, with agentic AI becoming neurosymbolic once code execution is paired with symbolic control.
  • Garcez and Lamb [2023] Artur d’Avila Garcez and Luís C. Lamb. Neurosymbolic AI: The 3rd wave. Artificial Intelligence Review, 56:12387–12406, 2023.
  • Gelman and Yao [2021] Andrew Gelman and Yuling Yao. Holes in Bayesian statistics. Journal of Physics G: Nuclear and Particle Physics, 48(1):014002, 2021. doi: 10.1088/1361-6471/abc3a5. Catalogues structural tensions in Bayesian inference, including the “Cantor’s corner” argument (hole 6): checking a model against data—necessary in practice whenever the model class is not known a priori to contain the truth—is not itself a coherent Bayesian operation, since it requires stepping outside the assumed model to entertain alternatives not covered by the prior.
  • Harnad [1990] Stevan Harnad. The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3):335–346, 1990. doi: 10.1016/0167-2789(90)90087-6. Canonical statement of the symbol-grounding problem we argue does not apply to our notion of grounding.
  • Hersbach et al. [2020] Hans Hersbach et al. The ERA5 global reanalysis. Quarterly Journal of the Royal Meteorological Society, 146(730):1999–2049, 2020. Single coherent best-estimate of the global atmospheric state.
  • Hoeting et al. [1999] Jennifer A. Hoeting, David Madigan, Adrian E. Raftery, and Chris T. Volinsky. Bayesian model averaging: A tutorial. Statistical Science, 14(4):382–417, 1999.
  • Huh et al. [2024] Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. The platonic representation hypothesis. In International Conference on Machine Learning, 2024.
  • Hutter [2005] Marcus Hutter. Universal Artificial Intelligence: Sequential Decisions Based on Algorithmic Probability. Springer, 2005. Defines AIXI, the unbounded Bayesian-optimal predictor/agent.
  • Kalnay [2003] Eugenia Kalnay. Atmospheric Modeling, Data Assimilation and Predictability. Cambridge University Press, 2003. Standard reference treating data assimilation (4D-Var, ensemble Kalman filtering) as Bayesian conditioning of a physical model.
  • Karniadakis et al. [2021] George Em Karniadakis, Ioannis G. Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang. Physics-informed machine learning. Nature Reviews Physics, 3(6):422–440, 2021. Embeds known governing equations into neural-network training as soft constraints; the canonical mechanism-regularized (but still neural, unverified) alternative to an explicit program.
  • Kaufmann et al. [2025] Rafael Kaufmann, Felix Neubürger, Michael Walters, Thomas Kopinski, and Dimitrije Marković. The CRISTAL method: Fast, reliable analytical problem-solving with pre-synthesized grounded world models. In Proceedings of the 19th Conference on Neurosymbolic Learning and Reasoning (NeSy), Proceedings of Machine Learning Research, 2025. Primordia / GAIA Lab. Introduces grounded world models (GWMs): a synthesized, continually-refined probabilistic program enabling full Bayesian inference; reaches Bayes-optimal accuracy on a synthetic-equities benchmark with 5 examples where SOTA LLMs plateau near 40%.
  • Kaufmann et al. [2026] Rafael Kaufmann, Harald Stromfelt, Thomas Minter, and Sandeep Ramesh. Coherent world-model meshes: Cross-case joint inference for structural causal prediction. Companion paper (Paper 2). Develops the cross-case Case Mesh: a shared-latent joint posterior across coupled cases, the cost of brute-forcing cross-case coherence on a scale-free collision graph, and the decision-impact benchmark for neglected downside correlation. Builds on and cites the present paper., 2026.
  • Koller and Friedman [2009] Daphne Koller and Nir Friedman. Probabilistic graphical models: Principles and techniques. MIT Press, 2009.
  • Koza [1992] John R. Koza. Genetic programming: On the programming of computers by means of natural selection. MIT Press, 1992. Founding text of genetic programming: evolving computer programs against a fitness measure rather than hand-writing them.
  • Kumar et al. [2025] Akarsh Kumar, Jeff Clune, Joel Lehman, and Kenneth O. Stanley. Questioning representational optimism in deep learning: The fractured entangled representation hypothesis. arXiv preprint arXiv:2505.11581, 2025.
  • Lam et al. [2023] Remi Lam et al. Learning skillful medium-range global weather forecasting. Science, 382(6677):1416–1421, 2023. GraphCast: ML weather emulator trained on ERA5 reanalysis.
  • LeCun [2022] Yann LeCun. A path towards autonomous machine intelligence. Open Review preprint, 2022.
  • Lewis et al. [2020] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 2020. Introduces retrieval-augmented generation (RAG): a non-parametric retriever supplies passages that a generator conditions on, the canonical long-context-retrieval augmentation of an LLM.
  • Liu et al. [2024] Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2024. doi: 10.1162/tacl_a_00638. Shows LLM performance on long-context tasks degrades significantly, and non-monotonically with position, as input context grows, even for models explicitly built for long contexts.
  • Lorenz [1963] Edward N. Lorenz. Deterministic nonperiodic flow. Journal of the Atmospheric Sciences, 20(2):130–141, 1963. Sensitive dependence on initial conditions; origin of the finite predictability horizon.
  • Lundberg and Lee [2017] Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, pages 4765–4774, 2017. SHAP: Shapley-value attributions for individual predictions.
  • Meyerson et al. [2025] Elliot Meyerson, Giuseppe Paolo, Roberto Dailey, Hormoz Shahrzad, Olivier Francon, Conor F. Hayes, Xin Qiu, Babak Hodjat, and Risto Miikkulainen. Solving a million-step LLM task with zero errors, 2025. Quantifies the cost of per-step error correction in long-horizon agentic execution: since per-step error rates do not vanish with scale, expected cost to complete an s-step task grows as Θ(slns) absent decomposition and voting-based verification, motivating extreme task decomposition to make error correction and rework tractable.
  • Mlodozeniec et al. [2025] Bruno Kacper Mlodozeniec, David Krueger, and Richard E. Turner. Position: Probabilistic modelling is sufficient for causal inference. In Proceedings of the 42nd International Conference on Machine Learning (ICML), volume 267 of Proceedings of Machine Learning Research, pages 81810–81840, 2025. Argues any causal inference question can be answered within standard probabilistic modelling and inference, reinterpreting causal-specific tools (e.g. the do-operator) as emerging from probabilistic modelling on a suitably expanded model rather than requiring bespoke causal notation.
  • Packer et al. [2023] Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: Towards LLMs as operating systems. arXiv preprint arXiv:2310.08560, 2023. Virtual context management: a hierarchical, agent-managed memory store that pages facts in and out of the LLM’s context across calls, the canonical instance of an “agentic memory” store.
  • Pearl [2009] Judea Pearl. Causal inference in statistics: An overview. Statistics Surveys, 3:96–146, 2009.
  • Quine [1951] Willard Van Orman Quine. Two dogmas of empiricism. The Philosophical Review, 60(1):20–43, 1951. Source of confirmation holism: statements face experience only as a corporate body, not one by one, against which we read our posture that “observables” are pragmatically defined by a model’s context of applicability rather than by any privileged ontological status.
  • Ribeiro et al. [2016] Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. “why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144, 2016. LIME: local surrogate explanations of individual predictions.
  • Richens and Everitt [2024] Jonathan Richens and Tom Everitt. Robust agents learn causal world models. In Proceedings of the 12th International Conference on Learning Representations (ICLR), 2024. Proves any agent satisfying a regret bound under a large set of distributional (interventional) shifts must have learned an approximate causal model of the data-generating process, converging to the true causal model for optimal agents.
  • Ro et al. [2025] Yeonju Ro, Haoran Qiu, Íñigo Goiri, Rodrigo Fonseca, Ricardo Bianchini, Aditya Akella, Zhangyang Wang, Mattan Erez, and Esha Choukse. Sherlock: Reliable and efficient agentic workflow execution. 2025. Documents that errors in agentic workflows propagate and compound across downstream steps, and quantifies the latency/cost overhead of the verification and rollback (rework) required to catch them, motivating selective, cost-optimal verification.
  • Schick et al. [2023] Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 2023. Self-supervised training of an LLM to decide when to call external tools (a calculator, Q&A system, search engine) and incorporate their outputs into generation.
  • Schmidt and Lipson [2009] Michael Schmidt and Hod Lipson. Distilling free-form natural laws from experimental data. Science, 324(5923):81–85, 2009. doi: 10.1126/science.1165893. Symbolic regression: searches jointly over equation form and parameters to recover free-form analytical laws from experimental data.
  • Sobieski and Biecek [2024] Bartosz Sobieski and Przemysław Biecek. Global counterfactual directions. In Proceedings of the European Conference on Computer Vision (ECCV), 2024. Discovers latent directions that flip a classifier’s decision across an entire dataset—a global, model-level counterfactual extraction, in contrast to local post-hoc attributions such as LIME and SHAP.
  • Solomonoff [1964] Ray J. Solomonoff. A formal theory of inductive inference, parts i and ii. Information and Control, 7(1–2):1–22, 224–254, 1964.
  • Spirtes et al. [2000] Peter Spirtes, Clark N. Glymour, and Richard Scheines. Causation, Prediction, and Search. MIT Press, 2nd edition, 2000. Canonical reference for constraint-based causal discovery (the PC/FCI algorithm family), which learns causal structure from observational data rather than from expert-informed priors.
  • The Decoder [2026] The Decoder. Claude sonnet 5 continues Anthropic’s pattern of hiding price increases behind unchanged token rates. https://the-decoder.com/claude-sonnet-5-continues-anthropics-pattern-of-hiding-price-increases-behind-unchanged-token-rates/, 2026. Claude Sonnet 5, released July 1, 2026; Artificial Analysis Intelligence Index score of 53 (max effort); standard pricing $3/$15 per 1M input/output tokens.
  • Vals AI [2026] Vals AI. Finance agent v2: Evaluating agents on core financial analyst tasks. https://www.vals.ai/benchmarks/fabv2, 2026. Benchmark of LLM agents on entry-level financial-analyst tasks over public filings; no model clears 58% and Claude Opus 4.8 scores 54%. The harness withholds code-execution and structured-memory tools.
  • Walters et al. [2025a] Michael Walters, Rafael Kaufmann, Justice Sefas, and Thomas Kopinski. Free energy risk metrics for systemically safe AI: Gatekeeping multi-agent study, 2025a. Primordia / GAIA Lab. Introduces a Cumulative Risk Exposure metric grounded in the Free Energy Principle for online, uncertainty-aware, preference-based risk governance in agentic and multi-agent systems, requiring only stakeholder-specified outcome preferences rather than exhaustive world models.
  • Walters et al. [2025b] Michael Walters, Thomas Kopinski, Rohil Rao, Rafael Kaufmann, and Alf Köhn-Seemann. The GAIA tech tree: A neurosymbolic AI framework for strategic technological decision-making. Working Paper, 2025b.
  • Wong et al. [2023] Lionel Wong, Gabriel Grand, Alexander K. Lew, Noah D. Goodman, Vikash K. Mansinghka, Jacob Andreas, and Joshua B. Tenenbaum. From word models to world models: Translating from natural language to the probabilistic language of thought, 2023. LLM-driven translation of natural language into probabilistic programs.
  • Yao et al. [2022] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022. Interleaves verbal reasoning traces with tool/environment actions (e.g. search, code execution) in a single LLM agent loop.