Extension SDK¶
OpenWAM extensions are ordinary installed Python modules. They register role-based components without modifying the OpenWAM source tree.
Compatibility Boundary¶
New integrations should import from these role-specific modules:
open_wam.sdk.config: typed config envelopes, loading, and resource resolutionopen_wam.sdk.data: sample contracts and dataset registrationopen_wam.sdk.policy: policy, decoder, attention, and visual runtime contractsopen_wam.sdk.simulator: simulator protocol and factory registrationopen_wam.sdk.results: versioned results and provenance
These modules are the compatibility-managed Python SDK. Other package exports
are not a promise that every implementation helper is stable. Modules below
open_wam.models.* are internal unless a contract is re-exported by
open_wam.sdk.policy.
The wheel ships a py.typed marker, so type checkers consume annotations from
these SDK modules directly. Public stability still follows the role-specific
SDK boundary above; typing visibility does not make internal implementation
modules compatibility-managed.
open_wam.sdk.config.resolve_experiment_config materializes fields owned by
typed config contracts. load_experiment_config and TrainingRuntime also
validate cross-section runtime and data requirements; use those boundaries
when starting an experiment rather than treating resolution as validation.
The base install supports config and result tooling. Install openwam[torch]
for dataset, policy, decoder, and attention extensions; use
openwam[train], openwam[eval], or openwam[sim] for the corresponding
runnable command. An extension package should declare the narrowest extra its
runtime actually needs.
Loading Extensions¶
Every maintained runtime command accepts repeatable extension specs:
openwam-train \
--extension acme_open_wam \
--extension research_runtime.install:register \
--cfg experiment.yaml
An extension spec is module[:hook]; the default hook is
register_open_wam. Hooks run once per process and in command-line order.
Import or registration failures stop startup before configuration is
constructed.
from open_wam.sdk.data import register_dataset_adapter
from .dataset import build_train_val
def register_open_wam() -> None:
register_dataset_adapter(
"acme_robot",
raw_builder=build_train_val,
description="ACME robot demonstrations.",
)
The same --extension contract is available on openwam-eval,
openwam-sanity, and openwam-sim-rollout. The module must be installed in
the active environment or otherwise importable on PYTHONPATH.
The packaged templates/extension_method/ example is runnable and imports
OpenWAM only through these SDK modules. Its module and config references are
the same from a source checkout and an installed wheel:
uv run --extra train openwam-train \
--cfg templates/extension_method/config.yaml \
--extension open_wam.templates.extension_method \
--save-root runs/extension-method-smoke
Its normalized visual-token policy and masked-MSE decoder are deliberately
small. Use them to verify registration, gradients, inference state, and
packaging before replacing one component at a time. The source scaffold lives
at src/open_wam/templates/extension_method; copy it into an
application-owned installable package before customization. Do not edit the
copy inside an installed OpenWAM wheel.
open_wam.templates.* is the namespace for runnable, copyable scaffolds; it is
not the extension API. Extension implementations depend on open_wam.sdk.*
and expose their own installed module name through --extension.
Registration APIs¶
- dataset adapters:
open_wam.sdk.data.register_dataset_adapter - policy variants:
open_wam.sdk.policy.register_policy_variant - action decoders:
open_wam.sdk.policy.register_action_decoder - simulator adapters:
open_wam.sdk.simulator.register_simulator_adapter
The active architectural boundary remains:
ExperimentConfig -> VariantPipeline -> VisualTower -> PolicyVariant -> ActionDecoder
Choosing The Extension Level¶
Use the first level that can express the change:
- Config only: select existing programs, layouts, schedules, cache policies, samplers, or decoder behavior in YAML.
- SDK extension: register an application-owned dataset adapter, policy variant, action decoder, or simulator adapter from an installed package.
- Core contribution: add a shared visual backend, exact sequence family, cache representation, training backend, or other reusable runtime contract.
Do not create a new architecture for a mask, loss, or data-layout change. Do not monkeypatch the shared runtime when the behavior needs a typed in-tree contract.
| Desired change | Owning boundary | Supported path |
|---|---|---|
| Change an existing program, geometry, loss weight, or optimizer setting | ExperimentConfig |
YAML or --set; no Python required |
| Read a new storage format, camera schema, or action/state representation | dataset adapter | Register raw_builder and/or latent_builder through open_wam.sdk.data |
| Supply real or counterfactual target-only FDM/IDM data | encoded dynamics artifact | Produce open_wam.encoded_dynamics.v1 and validate it with load_encoded_dynamics_artifact; routing stays dataset-agnostic |
| Check dataset-owned files before startup | dataset adapter | Register an artifact_resolver |
| Change dataset mixing, weighting, or distributed sample order | dataset and sampler | Implement the dataset sampling contract; shared advanced samplers remain provisional infrastructure |
| Add policy parameters, dense sequence semantics, or recurrent state | PolicyVariant |
Register an extension policy through open_wam.sdk.policy |
| Add dense attention visibility over the shared core | PolicyVariant and attention profile |
Submit a PreparedAttentionProfile through the dense runtime program |
| Change final action outputs, losses, sampling, or committed action count | ActionDecoder |
Register an extension decoder through open_wam.sdk.policy |
| Add a simulator backend | SimulatorBackend |
Register a factory through open_wam.sdk.simulator; the standard simulator CLI resolves the registered benchmark |
| Add an exact packed sequence family or cache tensor representation | VisualTower runtime |
Contribute a generic in-tree contract and parity tests; there is no runtime-backend registry |
| Replace the visual frontend, backbone, or decode stack | VisualTower |
Contribute in-tree and preserve checkpoint contracts |
| Add an optimizer, strategy, loop policy, checkpoint format, or log sink | training infrastructure | Select built-ins by config; new reusable implementations are currently in-tree contributions |
| Add an offline metric or application report | application evaluator | Build an application command around typed pipeline outputs |
For one encoded-dynamics view, use
open_wam.data.EncodedDynamicsLatentDataset.from_root(...). Orchestrators that
build several real/counterfactual or train/validation views should load one
open_wam.data.EncodedDynamicsResources per unique root and pass that shared
resource to each view; this avoids reparsing the complete transition index.
Core Boundary Ownership¶
ExperimentConfig¶
Start from the nearest maintained YAML and change one ownership axis at a
time. Shared finite choices stay in typed config fields. Dataset-specific
settings belong in data.adapter_options; policy- and decoder-specific
settings belong in their extension options mappings and should be parsed
into frozen application dataclasses by the extension builder.
Do not subclass ExperimentConfig in an extension package: the built-in YAML
loader will not discover that subclass. A new shared finite choice requires an
in-tree enum and named cross-section validation. Application-specific choices
remain in the open extension envelope.
Dataset adapters own source parsing, but strict FDM/IDM consumes a canonical
artifact rather than a benchmark-specific dataset class. Emit a manifest,
metadata/encoded_transitions.jsonl, model-space target latents, and aligned
action/state payloads following open_wam.encoded_dynamics.v1. The public
open_wam.sdk.data.load_encoded_dynamics_artifact validator checks the index;
positive route sources are preflighted before model construction.
VariantPipeline¶
VariantPipeline is composition infrastructure, not a plugin slot. It owns the
common train/inference order and connects one preprocessor, VisualTower,
PolicyVariant, and ActionDecoder. There is no
register_variant_pipeline API.
Customize through the adjacent contracts: request visual work from the policy, prepare policy inputs and decoder artifacts there, and implement final outputs and losses in the decoder. If those contracts cannot express a reusable behavior, extend the pipeline contract in-tree and add train and recurrent inference coverage for every maintained architecture.
VisualTower¶
The tower owns the shared frontend, visual core, runtime execution, and cache
lifecycle. Policies always receive frontend output and may additionally
request PolicyVisualStage.CORE from required_visual_stages().
Use the dense runtime plus PreparedAttentionProfile for custom visibility.
Adding a RuntimeProgramSpec name does not register an executor: a new exact
packing format, cache representation, or visual backbone requires an in-tree
runtime implementation and checkpoint/parity coverage. There is no
register_visual_tower API.
PolicyVariant¶
Use a policy extension for application-owned learned parameters,
packing/conditioning semantics, runtime-program selection, or recurrent
inference state. Declare standard shared-tower adapters through the extension
config's proprio_context_mode, dynamics_mode_context_enabled, and
text_conditioning_mode fields; they must be known before modules are
allocated. Return PolicyPipelineRequirements from pipeline_requirements()
to validate model-space geometry and carry any source-to-model action channel
projection into the decoder. Pass policy-specific outputs through a typed
DecoderArtifactEnvelope; do not expose decoder-private tensors through
unstructured pipeline keys. See
Adding A Policy Variant.
ActionDecoder¶
Use a decoder extension when policy topology remains valid but final
prediction, supervision, sampling, or rollout commitment changes. Keep
simulator-space conversion in the simulator adapter. A decoder that emits
routed-dynamics metrics should override dynamics_metric_namespace; auxiliary
validation reads that declared namespace without identifying the policy
architecture. A decoder can consume policy-declared assembly metadata through
configure_pipeline_requirements(). See Adding An Action Decoder.
Adding A Dataset¶
- Implement a dataset adapter that returns
WAMSample. - Keep source-specific parsing inside the adapter.
- Build canonical RGB layout in the data layer.
- Register raw and/or latent builders under one stable
dataset_type. - Add a focused adapter test and one config-loader smoke.
Use data.adapter_options for source-specific, open-ended settings. Keep
camera layout, action/state schema, sampling, and other shared semantics in
their typed data fields.
An adapter may expose:
raw_builder: returns raw-RGBWAMSampledatasetslatent_builder: returns pre-encodedLatentWAMSampledatasetsartifact_resolver: declares required files and directories for startup preflight- both builders under one key when a source supports both paths
Use DatasetArtifactRequirement instead of putting benchmark-specific path
checks in the trainer. The runtime evaluates these requirements before model
construction and reports the config owner and remediation for every missing
required artifact. Optional requirements are retained in run-start provenance.
from open_wam.sdk.data import DatasetArtifactKind, DatasetArtifactRequirement
def resolve_artifacts(config):
return (
DatasetArtifactRequirement(
name="episode manifest",
path=config.local_root,
kind=DatasetArtifactKind.DIRECTORY,
required=True,
config_path="data.local_root",
purpose="the ACME adapter discovers episodes below this root",
),
)
Registration rejects accidental replacement. Use a globally unique
dataset_type; replace=True is reserved for intentional process-local
overrides.
Distributed Sampling¶
A training dataset may implement
build_train_sampler(*, world_size, rank). Prefer the shared contracts in
open_wam.data when an adapter needs advanced distributed sampling. Those
sampler implementations are provisional infrastructure rather than part of
the narrow extension SDK:
WeightedReplacementDistributedSamplerfor seeded weighted drawsEpochOrderDistributedSamplerfor a dataset-provided global order padded to equal rank lengths; setgeometry_from_order=Truewhen weighting changes the epoch lengthEpochOffsetDistributedSamplerfor deterministic draw keys interpreted by the datasetPaddedEpochOffsetDistributedSamplerwhen draw-key epochs must not overlap after equal-rank paddingUnpaddedEpochOrderDistributedSampleronly when the training strategy explicitly supports unequal rank lengths
Construct one deterministic global order, then shard it by rank. Dataset adapters should own source weights and index interpretation, while these samplers own distributed coordination.
Policy And Decoder Envelopes¶
Application-owned policies and decoders use a typed outer envelope. The extension identifier is intentionally an open string; all built-in finite choices remain enums.
policy_variant:
name: extension
extension_type: acme.block_sparse_policy
hidden_size: 1536
attach_site: within_visual_core
text_conditioning_mode: task_prompt
options:
block_size: 64
history_chunks: 8
action_decoder:
name: extension
extension_type: acme.flow_action_decoder
hidden_size: 1536
action_dim: 7
action_horizon: 16
options:
loss: smooth_l1
options is copied into the frozen config envelope and must be a mapping with
string keys. Parse it into an application-owned typed dataclass inside the
builder. Shared data, backbone, training, and inference semantics stay in their
normal typed config sections.
Adding A Policy Variant¶
- Implement the
PolicyVariantcontract in the extension package. - Define its required methods:
required_visual_stages,prepare_train_inputs,forward_train,prepare_infer_state, andforward_infer_step. - Register a builder under the YAML
extension_type. - Add config, construction, gradient, and recurrent-inference tests.
A policy that participates in separate-model composition should declare its
artifact inputs and outputs in
PolicyInferenceCapabilities.composition_capabilities. Producers publish a
typed artifact such as PolicyGeneratedVideo; consumers receive the artifact
through an ordinary inference context and decide which available prompt,
proprio, and recurrent-history signals are visible. Keep model-specific
adaptation inside the policy variant by overriding resolve_inference_context.
The generic pipeline validates the neutral request once, then passes the
resolved context to both state preparation and inference. It must not transport
private caches or opaque in-flight state between models.
If decomposition must preserve a native stochastic program exactly, declare a
PolicyCompositionRngPolicy on the capability. caller_stream continues the
producer stream, including across CUDA devices; isolated_step_seed gives the
consumer an independent per-step stream. See
Video/Action Composition for the complete
contract and parity gates.
from open_wam.sdk.config import ExtensionPolicyConfig
from open_wam.sdk.policy import register_policy_variant
from .policy import AcmePolicy, AcmePolicyOptions
def build_policy(experiment):
config = experiment.policy_variant
assert isinstance(config, ExtensionPolicyConfig)
options = AcmePolicyOptions.from_mapping(config.options)
return AcmePolicy(config=config, options=options)
def register_open_wam() -> None:
register_policy_variant(
"acme.block_sparse_policy",
build_policy,
description="ACME block-sparse policy.",
)
Adding An Action Decoder¶
- Implement the
ActionDecodercontract in the extension package. - Keep final output construction, supervised loss, and action sampling in the decoder.
- Register a builder under the YAML
extension_type. - Add loss/output shape tests.
- Keep result schemas backward compatible when adding outputs.
from open_wam.sdk.config import ExtensionActionDecoderConfig
from open_wam.sdk.policy import register_action_decoder
from .decoder import AcmeActionDecoder, AcmeDecoderOptions
def build_decoder(experiment):
config = experiment.action_decoder
assert isinstance(config, ExtensionActionDecoderConfig)
options = AcmeDecoderOptions.from_mapping(config.options)
return AcmeActionDecoder(config=config, options=options)
def register_open_wam() -> None:
register_action_decoder("acme.flow_action_decoder", build_decoder)
The policy and decoder snippets are standalone examples. When one module owns
both, call both registration functions from the same register_open_wam hook.
Adding A Simulator¶
Simulator extensions register a factory under an application-owned benchmark identifier. The factory receives only immutable generic options and local-path aliases; it does not depend on OpenWAM's CLI parser.
from open_wam.sdk.simulator import (
SimulatorFactoryContext,
register_simulator_adapter,
)
from .simulator import AcmeSimulator
def build_simulator(context: SimulatorFactoryContext) -> AcmeSimulator:
return AcmeSimulator(
endpoint=context.options["endpoint"],
asset_root=context.local_paths.get("simulators.acme_root"),
)
def register_open_wam() -> None:
register_simulator_adapter("acme", build_simulator)
Invoke it without changing the OpenWAM repository:
openwam-sim-rollout \
--extension acme_open_wam \
--benchmark acme \
--sim-option endpoint=localhost:5000 \
--cfg experiment.yaml
The returned object implements SimulatorBackend. Task, episode, and seed are
passed through EpisodeSpec at reset time; application-specific construction
values belong in repeatable --sim-option KEY=VALUE settings.
Custom Attention¶
Attention visibility is data passed through the policy/runtime boundary, not a
backbone subclass. A custom policy can build a
PreparedAttentionProfile and pass it in
VisualCoreInput.attention_profile through the dense runtime program:
from open_wam.sdk.policy import (
PreparedAttentionProfile,
RuntimeStepInput,
VisualCoreInput,
build_dense_runtime_program,
)
result = visual_tower.execute_runtime_step(
RuntimeStepInput(
program=build_dense_runtime_program(),
core_input=VisualCoreInput(
tokens=tokens,
attention_profile=prepared_profile,
),
)
)
The profile may provide dense boolean masks, FlexAttention block masks, or both. Keep sequence packing and mask construction in the policy extension; the shared visual runtime applies backend-ready profiles and owns backbone execution. The exact parallel-stream and dual-expert backends are checkpoint compatibility contracts with fixed layout semantics, not general attention extension points.
The following built-in role modules are useful when contributing to OpenWAM itself, but are not stable extension APIs. The built-in attention implementation has three parameter-free roles:
open_wam.models.common.attention_contractsowns profile records, coupling names, and semantic normalization;open_wam.models.common.chunked_attentionassembles the maintained packed chunk masks; andopen_wam.models.common.attention_backendsselects and executes dense SDPA or FlexAttention representations.
Use the role module matching the operation when working on these implementations.
Built-in DualExpert checkpoint layouts are narrower policy-internal contracts:
open_wam.models.policy_variants.dual_expert.attention_unpackedowns dense diagnostic and checkpoint-compatibility layout builders; it is not a maintained training or GJD execution path;open_wam.models.policy_variants.dual_expert.attention_packedowns the exact packed coupling profiles used by maintained training; andopen_wam.models.policy_variants.dual_expert.attention_cachedowns split-cache action inference layouts.
Extensions implementing a new attention paradigm should normally
construct a common PreparedAttentionProfile; depend on a DualExpert role only when
the extension deliberately implements that exact built-in sequence layout.
Built-in DualExpert runtime controls are also split by parameter-free role:
open_wam.models.policy_variants.dual_expert.runtime_routesselects a typed runtime route;open_wam.models.policy_variants.dual_expert.rollout_geometryresolves chunk, history, cache, and action-execution geometry;open_wam.models.policy_variants.dual_expert.coupling_semanticsresolves block and timestep coupling; andopen_wam.models.policy_variants.dual_expert.inference_backendvalidates and restores the inference backend selected by a route.
These modules document the fixed built-in checkpoint contract; they are not a
registration API. A custom policy should express its behavior through its
PolicyVariant, runtime program, prepared attention profile, and decoder
rather than adding architecture-specific branches to these Dual Expert owners.
Shared Transformer Primitives¶
open_wam.models.visual_tower exports the shared Wan-style transformer
building blocks used by the visual core and action-side experts. Their
canonical owners are:
shared_transformer_supportforSharedTransformerAttentionandSharedTransformerBlock, including the stable attention-backend patch point;shared_transformer_embeddingsfor timestep and rotary positional embeddings plus rotary application;shared_transformer_layoutfor chunk-slice and split-segment tensor helpers;runtime_parameter_opsfor FSDP-safe linear, normalization, and feed-forward helpers.
Implementation code should import the role owner it uses. These functions own
learned transformer execution or the explicit tensor
operations supporting it, not sequence visibility or cache retention.
Extensions should normally submit an attention profile through
VisualCoreInput; use the lower-level primitives only when implementing a
genuinely new reusable block outside the built-in core.
Cache Policy¶
The parameter-free cache API is split by role:
open_wam.models.common.cache_backend_contractsdefines backend specs, payload records, and backend selection;open_wam.models.common.cache_layout_policydefines attention-mask, prefix-visibility, packed-sequence, slot-retention, and prefix-merge policy;open_wam.models.common.cache_backend_lifecycledefines payload allocation, mutation, reset, and materialization.
Custom policies that retain the built-in cache formats can reuse:
prepare_sdpa_maskandprepend_cached_prefix_mask;resolve_slot_pool_prefix_visibility;packed_slot_pool_query_sequence_ids;retained_slot_pool_indices_for_current_write;merge_attention_cache_entries.
Custom runtime programs can use init_cache_backend_payload,
update_slot_pool_layer_state, clear_cache_backend_payload, and
materialize_cache_backend_entries from cache_backend_lifecycle without
depending on transformer execution. The payload types come from
cache_backend_contracts; do not duplicate their tensor-layout conventions in
a policy variant.
open_wam.models.visual_tower.RuntimeCacheLifecycle composes those backend
operations into initialization, named-branch, retention, cursor-advance, and
reset operations over the public CacheState contract. VisualTower exposes
the same operations as its stable runtime facade and supplies the current
backbone capability and layer count dynamically. The lifecycle is a frozen
plain object, not an nn.Module, so using or replacing it cannot add
checkpoint keys.
Runtime-backbone checkpoint selection and operational compatibility live in
open_wam.models.visual_tower.runtime_backbone. These helpers borrow the
tower-owned module rather than wrapping it, so custom runtime programs can
reuse access validation, device normalization, and cache reset without
creating a second parameter owner.
Attention profiles decide which tokens may interact. Cache policy decides how already-computed keys and values are represented, retained, and prepended. Neither contract owns learned parameters. A new cache representation still requires a backend integration; do not encode its retention rules inside a policy variant or transformer block.
Training And Checkpoints¶
Built-in optimizer, scheduler, precision, strategy, loop, checkpoint, and
logging choices are config driven. They are not SDK registries. An application
may call a VariantPipeline from its own research loop, but then it owns
distributed coordination, resume semantics, validation, and logging. Reusable
behavior belongs in a generic typed in-tree training contract, not a policy- or
benchmark-named trainer branch.
Registered policy and decoder modules participate in normal model and full-training-state checkpoints. Treat parameter names, module registration order, tensor shapes, optimizer mapping, and recurrent cache semantics as compatibility contracts. A topology change creates a new checkpoint contract; an orchestration-only change must preserve output, loss, gradient, optimizer, and recurrent-inference parity.
Extension Workflow¶
- Copy the nearest maintained config and identify the owning boundary.
- Keep application code in an installable package outside OpenWAM's built-in implementation directories.
- Parse every open
optionsmapping into a frozen application dataclass. - Implement the smallest supported contract and register it from one side-effect-free hook.
- Pass the same
--extension module[:hook]to every train, eval, sanity, and rollout command that constructs the component. - Test config parsing, shapes and masks, one forward/backward update, checkpoint load, and recurrent inference. Add distributed and simulator tiers when the component uses them.
- Retain the resolved config and artifact manifest for every reported run.
Contract Rules¶
- Registration hooks configure contracts; they must not start jobs or mutate global training state.
- Use globally unique extension identifiers. Duplicate registration fails unless the caller explicitly requests a process-local replacement.
- Use
open_wam.sdk.config.coerce_fieldsfor enum-backed extension settings; keep cross-section defaults and validation in a named config contract. - Dataset parsing remains in data adapters.
- Policy semantics remain in
PolicyVariant. - Shared visual execution remains in
VisualTower. - Final supervised outputs and losses remain in
ActionDecoder. - Public finite choices are enum-backed; dataset names, row keys, paths, and extension labels remain open strings.
Cookbooks¶
docs/cookbooks/new_policy_architecture.mddocs/cookbooks/new_action_decoder.mddocs/cookbooks/new_dataset.mddocs/cookbooks/new_simulator_adapter.mddocs/cookbooks/reproduce_result.md