Skip to content

Repository files navigation

Expected free energy as an information constraint on the Bethe Lagrangian

Code and paper source for:

Wouter M. Kouw. Expected free energy as an information constraint on the Bethe Lagrangian. International Workshop on Active Inference (IWAI), 2026.

Active inference is formulated here as minimisation of a constrained Bethe free energy functional. Instead of adopting expected free energy (EFE) as the planning objective, the epistemic drive is imposed as an explicit information constraint: the mutual information between predicted observations and the latent states and parameters, averaged over the action posterior, must be at least the entropy of the goal prior.

Minimising the constrained Bethe Lagrangian under normalisation, marginalisation, form and information constraints gives the stationary policy

q*(u_τ) ∝ p(u_τ) · exp( Σ_t E_{q(y_t|u_τ)}[ln p_*(y_t)] + γ_t · I[x_t, φ ; y_t | u_τ] )

where the weight γ on the information term is a Karush–Kuhn–Tucker multiplier solved from the information floor, not a tuned hyperparameter. At γ = 1 the policy coincides exactly with the EFE policy q(u_τ) ∝ p(u_τ) exp(−G(u_τ)); at γ = 0 the constraint is inactive and the epistemic drive switches off.

Repository layout

Path Contents
agents/common.py Shared machinery. DiscreteBeliefAgent — model validation, HMM forward filtering, randomised argmax tie-break. PolicyEnumerationAgent — goal prior, joint emission/goal construction, enumeration of the `
agents/bethe_agent/ ConstrainedBetheAgent — the planner of the paper. Computes rollout features in one vectorised forward sweep and solves the scalar dual γ.
agents/efe_agent/ ExpectedFreeEnergyAgent — standard EFE baseline scoring policies by G(u_τ) = risk + ambiguity (− novelty in Dirichlet mode).
agents/qmdp_agent/ HMMPlanner — Q-MDP over a known discrete HMM, the exploit-only baseline.
agents/test_agents.py Smoke tests for the shared machinery and the γ = 1 equivalence (python agents/test_agents.py, ~1 s).
envs/ Discrete environments and their generative models: tmaze, cue_grid, gate_cue. Optional Navix/JAX gridworlds are lazily imported and unused by the paper's experiments.
experiment-Tmaze/ Canonical T-maze, including the γ = 1 equivalence check against EFE.
experiment-cuegrid/ Cue-grid navigation, plus the information-floor sweep.
experiment-gatecue/ Salience-vs-novelty allocation under learning.
figures/ Figures rendered into the paper.
graphs/ TikZ sources for the factor-graph figures.
main.tex, references.bib Paper source.

Installation

Python 3.10 or newer.

pip install numpy scipy casadi matplotlib

CasADi supplies the Newton rootfinder (with exact Jacobians) used for the dual; SciPy is used only for the digamma in the Dirichlet novelty term; Matplotlib is needed for the figure scripts. No installation step is required for the package itself — the experiment scripts insert the repository root on sys.path, so run everything from the repository root.

Reproducing the paper

All experiments run 200 trials from fixed seeds and write one .npz per agent plus a config.json into their own results/ directory. Those results are committed, so experiment-cuegrid/make_figure.py and experiment-gatecue/make_figure.py can be run without re-running the experiments. The floor-sweep figure is self-contained: it runs its own sweep rather than reading results/.

# Table 1 — canonical T-maze, and the γ = 1 ⇒ EFE equivalence check
python experiment-Tmaze/run_experiment.py

# Table 2 and Figure 2 — cue grid
python experiment-cuegrid/run_experiment.py
python experiment-cuegrid/make_figure.py

# Figure 3 — the information floor selects the dual's regime
python experiment-cuegrid/make_figure_floor_sweep.py

# Figure 4 — salience/novelty allocation under learning
python experiment-gatecue/run_experiment.py
python experiment-gatecue/make_figure.py

The cue-grid and gate-cue runners accept overrides such as --n-trials, --horizon and --max-steps; see --help. The T-maze runner and the floor probe have no command-line interface — adjust the constants at the top of the script instead.

Two supporting runs are not in the paper but document the claims around it:

# Does varying the goal-prior width traverse all three KKT regimes?
python experiment-cuegrid/run_floor_variance_probe.py

# Can standard EFE recover cue-seeking through loss-averse preference design
# rather than through the information weight?
python experiment-cuegrid/run_efe_lossaverse.py

Each experiment directory also holds a visualize.ipynb for inspecting results interactively.

For a quick check that the agents still behave after a change, python agents/test_agents.py runs in about a second. The authoritative check is re-running the experiments above and confirming results/ is unchanged — every array is deterministic under the fixed seeds except plan_time_per_step, which records wall-clock timings.

The equivalence check

experiment-Tmaze/run_experiment.py runs the constrained Bethe agent with gamma_override=1.0 alongside the EFE agent under a shared goal prior and compares policy posteriors on every planning call of all 200 trials. The maximum deviation is written to results/gamma1_equivalence.json and is at machine precision (≈ 2.2e-16), verifying Proposition 1 numerically.

Using the agent

from agents.bethe_agent import ConstrainedBetheAgent

agent = ConstrainedBetheAgent(
    A,                      # list of emission matrices, one per observation modality
    B,                      # transition tensor B[s', s, a]
    D,                      # initial state prior
    C,                      # factored goal prior p_*(y) per modality
    horizon=3,
    param_posterior="dirichlet",   # "point" for known transitions
    gamma_max=1e3,
)

agent.reset()
action = agent.step(obs)    # obs is one index per modality

After each step, the solved dual and its regime are exposed on agent.last_gamma and agent.last_gamma_status ("inactive", "root", "saturated", or "override"), alongside last_q_pi, last_q_u1 and last_info_per_step.

Two knobs matter most:

  • param_posterior. In "point" mode the transitions are taken as known and the information gain reduces to state salience I[x_t; y_t | φ, u_t]. In "dirichlet" mode an independent Dirichlet posterior is maintained per column of B, updated from observed transitions, and the parameter novelty I[φ; y_t | u_t] is added — the chain-rule split of the paper. The Dirichlet posterior deliberately survives reset(), so knowledge accumulates across trials; use reset_parameters() to clear it.
  • gamma_override. Left as None, γ is solved from the floor. Set it to 1.0 to recover standard EFE, or to any fixed weight to treat the epistemic bonus as a tuned precision. Only a scalar is accepted; the per-step γ_t vector was removed once the horizon constraint was aggregated.

How the dual is solved

The per-step information constraints are aggregated into a single horizon constraint with one shared multiplier γ, which restricts the per-step duals to the diagonal γ_t ≡ γ — the ray on which the EFE equivalence lives. Because the expected information gain is monotone in γ within the exponential family generated by the stationary policy, complementary slackness leaves exactly three cases:

  • the floor is already met at γ = 0 → inactive, no epistemic drive;
  • the binding condition has a unique interior root → root, located by Newton with a bracketed-bisection fallback on [0, γ_max];
  • the floor exceeds the information any rollout can supply → saturated at γ_max (10³ in all experiments), the maximal information-seeking limit.

Under the full-joint goal-prior floor of the benchmarks the third case is realised. Restricting the floor to the reward modality and varying its width exercises the other two — this is what Figure 3 sweeps.

Building the paper

latexmk -pdf main.tex

Camera-ready revisions are wrapped in a \revised{...} macro that typesets them in blue. Redefine it in the preamble of main.tex as \newcommand{\revised}[1]{#1} to render the final version in black.

Citation

@inproceedings{kouw2026expected,
  title     = {Expected free energy as an information constraint on the {Bethe} {Lagrangian}},
  author    = {Kouw, Wouter M.},
  booktitle = {International Workshop on Active Inference},
  year      = {2026}
}

About

Companion repository to a submission to the International Workshop on Active Inference 2026 focused on casting the expected free energy functional as a constraint on the Bethe free energy functional.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages