PPDL: Probabilistic PDL
PPDL is the probabilistic extension of PDL, introduced in the paper
PPDL: LLM-Based Flows as Probabilistic Programs
(ICML 2026).
It turns any PDL program into a probabilistic program by adding a single new construct — factor — and
a set of plug-and-play probabilistic inference engines.
Motivation
LLM-based flows accumulate uncertainty: each step depends on the (probabilistic) output of the previous one, so errors compound. PPDL addresses this in two complementary ways:
- Quantify uncertainty — instead of a single output value, the runtime returns a distribution over possible outputs, giving users a principled view of confidence.
- Improve accuracy — inference engines (importance sampling, SMC, majority voting) explore many execution traces in parallel and weight them by user-supplied constraints, steering the search toward more correct answers.
Because the inference engine is orthogonal to the program logic, you write the flow once and switch between algorithms at runtime without touching a single line of PDL.
The factor Block
PDL already provides the first core construct of probabilistic programming — sample — in the form of an LLM model call. PPDL adds the second: factor.
factor: <score_expression>
factor updates the score (log-probability) of the current execution trace by the value of the
expression. The expression must evaluate to a float.
- A positive value makes the current trace more likely (rewards).
- A negative value makes it less likely (soft penalty).
−∞(or a very large negative number like-1000) is a hard constraint — it effectively eliminates the trace.
The result of a factor block is the empty string "" and it does not contribute to the background
context.
Simple example: biased coin
defs:
obs:
array: [0, 0, 1, 1, 1, 1, 1, 1, 1, 1]
p:
model: ollama_chat/granite4:micro
parameters:
temperature: 1
input: Generate a number between 0 and 1 that represents the bias of a coin that might returns head (represented with the value 1) more often than tail (represented with the value 0). DO NOT GENERATE A PYTHON CODE, JUST ANSWER WITH THE NUMBER
parser: json
spec: number
fallback:
lastOf:
- factor: -1000
- data: 0.5
lastOf:
- defs:
score:
lang: python
code: |
import math
# Calculate log probability of observing the data given p
log_prob = 0.0
for flip in obs:
if flip == 1:
# Probability of heads
log_prob += math.log(p) if p > 0 else -1000
else:
# Probability of tails
log_prob += math.log(1 - p) if p < 1 else -1000
result = log_prob
factor: ${score}
- ${p}
Here the LLM is asked to guess the bias p of a coin given a sequence of observations. The
factor block scores the trace by the log-likelihood of the observed flips given p, so that
after inference the distribution of p values concentrates around the one most consistent with
the data.
Running PPDL Programs
PPDL programs are executed with the pdl-infer command-line tool (installed alongside pdl):
pdl-infer [OPTIONS] <program.pdl>
| Option | Default | Description |
|---|---|---|
--algo |
smc |
Inference algorithm (see below) |
-n, --num-particles |
5 |
Number of particles |
-w, --workers |
unlimited | Max parallel workers |
-d, --data |
— | Initial scope variables (YAML/JSON string) |
-f, --data-file |
— | YAML file containing initial scope variables |
-v, --viz |
off | Display bar chart of the output distribution |
The tool prints a sample drawn from the output distribution.
Example
# Run with Sequential Monte Carlo, 10 particles, 4 parallel workers
pdl-infer --algo smc -n 10 -w 4 examples/ppdl/coin.pdl
Inference Algorithms
PPDL provides four families of inference engines, each available in a sequential and a parallelized variant.
Importance Sampling (is / parallel-is)
Runs N independent full traces of the program and weights the results by their accumulated
score. Use this when the constraint information is mostly available at the end of a trace.
Sequential Monte Carlo (smc / parallel-smc) — default
Runs N particles and resamples after each factor block: particles with a low score are
replaced by copies of higher-scoring particles. This redirects compute toward more promising
partial traces during execution, making it particularly effective for deep multi-step flows
with informative intermediate constraints.
Resampling points correspond to factor blocks in the program.
Majority Voting (maj / parallel-maj)
Runs N independent traces and ignores all factor scores; the output distribution is uniform.
Use this as a simple ensemble baseline when you do not have meaningful factors.
Rejection Sampling (rejection / parallel-rejection)
Repeatedly runs the program until N accepted samples are collected. Accepts a trace with
probability proportional to its score.
Adding Constraints to a Flow
The real power of PPDL comes from combining a regular PDL flow with one or more factor blocks
that encode soft or hard constraints.
LLM-as-a-Judge
Use an LLM call to score intermediate or final results:
defs:
llm: ollama/granite4:micro
problem_statement: |
Write a function to find nth centered hexagonal number.
assert centered_hexagonal_number(10) == 271
lastOf:
- >
Generate an English plan for how to generate code for
the following problem: ${ problem_statement }
- model: ${ llm }
def: plan
# Score the plan using an LLM-as-judge
- defs:
score:
lastOf:
- Is this plan for the problem correct?
Problem: ${ problem_statement }
Plan: ${ plan }
- model: ${ llm }
parser: json
spec: boolean
# log-probability: 0 if true, -inf if false
factor: ${ 0 if score else -1000 }
- >
Generate a complete Python function definition for: ${ problem_statement }
- model: ${ llm }
def: solution
- ${ solution }
Note
The PPDL paper uses a continuous score based on log-probabilities of the judge's true/false
output tokens, which is available via the expectations sugar (see below).
Rule-Based Constraints
Hard or soft constraints derived from deterministic checks:
defs:
count_warnings:
function:
code: string
return:
lang: python
code: |
import subprocess
result = subprocess.run(
"flake8 --isolated -", input=code, capture_output=True,
shell=True, text=True, check=False
)
result = len(result.stdout.strip().splitlines())
...
# Geometric prior: penalise each additional warning
- defs:
warnings:
call: ${ count_warnings }
args: { code: ${ solution.code } }
import math
penalty: ${ math.log(0.5) * warnings }
factor: ${ penalty }
Hard Constraints
Set the factor to a very large negative number to eliminate a trace entirely:
factor: ${ 0 if valid else -1000 }
or use factor: -1000 directly for a constant hard penalty.
The expectations Sugar
For common LLM-as-a-judge patterns, PDL provides the expectations field on any block.
It is shorthand that automatically wires up an LLM judge and a factor:
model: ${ llm }
def: solution
expectations:
- expect: This solution is correct for the following problem.
# optional custom feedback function; defaults to stdlib LLM judge
The runtime evaluates each expectation using the configured judge model and applies the
resulting score as a factor. The feedback field allows providing a custom scoring
function that returns either a float score or a [float, str] pair where the string is
feedback injected into the context for retrying.
Python SDK
Running a PPDL program
from pdl.pdl_infer import exec_file, exec_str, PpdlConfig
from pdl.pdl_distributions import Categorical
ppdl_config = PpdlConfig(
algo="parallel-smc", # or "is", "maj", "rejection", ...
num_particles=10,
max_workers=4,
)
dist: Categorical = exec_file("examples/ppdl/coin.pdl", ppdl_config=ppdl_config)
# Inspect the distribution
dist_sorted = dist.sort() # sort by probability (highest first)
best = dist_sorted.values[0] # most probable output
prob = dist_sorted.probs[0] # its probability
print(f"Best answer: {best} (p={prob:.4f})")
# Or sample from the distribution
sample = dist.sample()
PpdlConfig reference
See PpdlConfig in the API reference.
exec_file / exec_str / exec_dict / exec_program
See the Probabilistic Inference section of the API reference.
Categorical distribution
See Categorical in the API reference.
Built-in Special Variables
During PPDL execution, the interpreter maintains two special scope variables:
| Variable | Description |
|---|---|
pdl_context |
The accumulated conversation context (list of role/content messages). |
pdl_score |
The running log-probability of the current trace, updated by each factor block. |
pdl_score is not typically accessed in PDL programs directly; it is managed by the runtime
and used by the inference engine for weighting and resampling.
The variable pdl_particle_id (an integer) is injected into the scope for each particle,
which is useful when you need different behavior per particle (e.g., different random seeds).
Parallel Execution
PPDL parallelizes execution at two levels:
-
Across particles — multiple traces are executed concurrently using a thread pool (
parallel-is,parallel-smc, etc.). For SMC, a synchronization point occurs at eachfactorblock for resampling. -
Within a single trace — independent model calls within one trace are parallelized using futures. The interpreter only awaits a model response when the result is actually needed, which is especially effective inside
defsblocks where several variables can be computed concurrently.
Control the number of worker threads with the -w / --workers CLI flag or the
max_workers key in PpdlConfig.
Complete Examples
Name suggestion with demographic priors (examples/ppdl/name_finder.pdl)
Demonstrates using factor to steer an LLM toward names that are popular in two different
geographic populations.
defs:
model: ollama/granite4:micro
names_fr:
data: {
"Louise": {"2020": 5.30, "2021": 5.33, "2022": 4.94, "2023": 4.69, "2024": null},
"Jade": {"2020": 5.30, "2021": 5.31, "2022": 4.95, "2023": 4.26, "2024": null},
"Ambre": {"2020": 3.82, "2021": 4.29, "2022": 4.89, "2023": 4.67, "2024": null},
"Alba": {"2020": 2.09, "2021": 3.55, "2022": 4.75, "2023": 4.55, "2024": null},
"Emma": {"2020": 4.84, "2021": 4.97, "2022": 4.57, "2023": 3.93, "2024": null},
"Gabriel": {"2020": 6.14, "2021": 6.39, "2022": 7.08, "2023": 6.68, "2024": null},
"Leo": {"2020": 6.25, "2021": 6.10, "2022": 5.91, "2023": 5.10, "2024": null},
"Raphael": {"2020": 5.52, "2021": 5.53, "2022": 5.50, "2023": 5.13, "2024": null},
"Mael": {"2020": 4.58, "2021": 4.83, "2022": 5.17, "2023": 4.84, "2024": null},
"Louis": {"2020": 5.28, "2021": 5.25, "2022": 5.16, "2023": 4.91, "2024": null},
"Remi": {"2020": 1.78, "2021": 1.70, "2022": 1.67, "2023": 1.62, "2024": null},
"Rosa": {"2020": 1.18, "2021": 1.28, "2022": 1.40, "2023": 1.51, "2024": null},
"Josephine": {"2020": 2.18, "2021": 2.34, "2022": 2.49, "2023": 2.65, "2024": null},
"Odilon": {"2020": 0.42, "2021": 0.45, "2022": 0.51, "2023": 0.59, "2024": null}
}
names_us:
data: {
"Gabriel": {"2020": 3.55, "2021": 3.41, "2022": 3.34, "2023": 3.26, "2024": 3.18},
"Louis": {"2020": 0.82, "2021": 0.79, "2022": 0.78, "2023": 0.77, "2024": 0.75},
"Jules": {"2020": 0.41, "2021": 0.40, "2022": 0.39, "2023": 0.38, "2024": 0.38},
"Arthur": {"2020": 0.68, "2021": 0.65, "2022": 0.64, "2023": 0.62, "2024": 0.61},
"Noah": {"2020": 5.18, "2021": 5.05, "2022": 5.01, "2023": 4.96, "2024": 4.92},
"Adam": {"2020": 2.73, "2021": 2.67, "2022": 2.67, "2023": 2.67, "2024": 2.67},
"Isaac": {"2020": 2.18, "2021": 2.13, "2022": 2.11, "2023": 2.10, "2024": 2.08},
"Jade": {"2020": 2.18, "2021": 2.13, "2022": 2.11, "2023": 2.10, "2024": 2.08},
"Louise": {"2020": 0.33, "2021": 0.31, "2022": 0.31, "2023": 0.30, "2024": 0.29},
"Emma": {"2020": 4.09, "2021": 3.95, "2022": 3.89, "2023": 3.83, "2024": 3.77},
"Rose": {"2020": 1.36, "2021": 1.34, "2022": 1.33, "2023": 1.33, "2024": 1.33},
"Alice": {"2020": 1.09, "2021": 1.06, "2022": 1.06, "2023": 1.05, "2024": 1.04},
"Anna": {"2020": 2.46, "2021": 2.40, "2022": 2.39, "2023": 2.38, "2024": 2.37},
"Lina": {"2020": 0.82, "2021": 0.79, "2022": 0.78, "2023": 0.77, "2024": 0.75},
"Remi": {"2020": 0.33, "2021": 0.31, "2022": 0.31, "2023": 0.30, "2024": 0.29},
"Rosa": {"2020": 0.25, "2021": 0.24, "2022": 0.24, "2023": 0.24, "2024": 0.24},
"Josephine": {"2020": 0.96, "2021": 0.93, "2022": 0.92, "2023": 0.91, "2024": 0.90},
"Odilon": {"2020": 0.01, "2021": 0.01, "2022": 0.02, "2023": 0.02, "2024": 0.02}
}
names:
lang: python
code: |
result = [ name for name in names_us.keys() ] + [ name for name in names_fr.keys() ]
is_rare:
function:
name: string
names: object
return:
lang: python
code: |
stats = names.get(name)
if stats is None:
result = -1000
else:
if stats["2023"] < 0.01:
result = -100
else:
result = -stats["2023"]
name:
lastOf:
- |
I am looking for a first name for my future baby. Can you give me one suggestion? You must to pick one from the following list:
${names}
- model: ${ model }
parameters:
temperature: 1.5
- Extract ONLY the name
- model: ${ model }
condition_one_word:
defs:
len:
lang: python
code: |
result = len(name.split())
if: ${ len > 1 }
then:
factor: -100
condition_american:
defs:
penalty:
call: ${is_rare}
args:
name: ${name}
names: ${names_us}
factor: ${penalty}
condition_french:
defs:
penalty:
call: ${is_rare}
args:
name: ${name}
names: ${names_fr}
factor: ${penalty}
text:
- ${name}
Hidden Markov Model posterior (examples/ppdl/hmm.pdl)
Shows PPDL as a general probabilistic programming language: a classical state-space model
where the LLM is replaced by a Python sampler, and factor encodes the observation
likelihood.
defs:
rng:
lang: python
code: |
import random
import numpy as np
seed_value = random.randint(1, 10000)
result = np.random.default_rng(seed_value)
step:
function:
pre_x: number
y: number
return:
defs:
x:
lang: python
code: |
from scipy.stats import norm
result = norm(pre_x, 1).rvs(random_state=rng)
score:
lang: python
code: |
from scipy.stats import norm
result = norm(x, 1).logpdf(y)
lastOf:
- factor: ${score}
- data: ${x}
obs:
array: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 25, 26, 27, 28, 29, 30]
pre_x: 0
for:
y: ${obs}
repeat:
defs:
pre_x:
call: ${step}
args:
pre_x: ${pre_x}
y: ${y}
data: ${pre_x}
join:
as: lastOf
Code generation with constraints (examples/ppdl/mbpp.pdl)
A MBPP benchmark program that combines an LLM-as-a-judge (expectations) with a
rule-based flake8 linter score.
Probabilistic ReAct agent (examples/prompt_library/gsm8k_prob_react.pdl)
Extends the standard ReAct agent with factor blocks that penalize invalid actions and
long trajectories, giving the SMC engine useful intermediate signals for resampling.
Relationship to PDL
PPDL is a strict superset of PDL:
- Every PDL program is a valid PPDL program. Running it with
pdl-infer -n 1is equivalent to running it withpdl. - The
factorblock is the only new language construct. - The
expectationsfield (available on any block) is syntactic sugar built on top offactorand the existingretrymechanism. - All PDL features — variables, conditionals, loops, functions, tool calls, parsers, type-checking, imports — work unchanged inside PPDL programs.
The key difference is the execution model: whereas pdl always runs a single trace and
returns one result, pdl-infer runs many traces, weights them, and returns a Categorical
distribution over results.
Further Reading
- PPDL paper (ICML 2026)
- PDL Tutorial — full PDL language reference
- API Reference — Python SDK documentation