Parallel constrained decoding for structured extraction and classification on Apple Silicon.
pcd-rlcd extracts the core inference engine from the original Parallel Constrained Decoding project and makes it available as a small, reusable Python package.
It evaluates multiple schema fields in parallel instead of generating a complete JSON document token by token. This is useful for classification, decision routing, structured extraction, and other workloads where fields have bounded candidate values.
- Parallel evaluation of multiple structured fields
- Enum and boolean schema fields
- Constrained candidate selection
- Programmatic JSON assembly
- Field-level confidence scores
- Top-choice probability telemetry
- MLX backend for Apple Silicon
- Torch backend support
- Configurable language model through
MODEL_ID - Typed Python package
- Benchmark utilities
- Python 3.10+
- Apple Silicon Mac for the MLX backend
- macOS 14 or later recommended
- A model supported by the selected inference backend
The package defaults to:
mlx-community//Qwen2.5-1.5B-Instruct
The example configuration uses:
mlx-community/Qwen2.5-0.5B-Instruct-4bit
You can select a different model with the MODEL_ID environment variable.
Using uv:
uv add pcd-rlcdUsing pip:
pip install pcd-rlcdTo install directly from the repository:
git clone <repository-url>
cd <repository-directory>
uv syncfrom pcd_rlcd.schema import StructuredSchema
from pcd_rlcd.engine import run_parallel_generation
schema_definition = {
"fraud_risk": {
"type": "enum",
"choices": ["LOW", "ELEVATED", "SUSPICIOUS", "CRITICAL"],
"description": "Risk assessment tier for incoming transaction",
},
"block_account": {
"type": "boolean",
"description": "Whether immediate account restriction is required",
},
"recommended_action": {
"type": "enum",
"choices": [
"ALLOW",
"STEP_UP_2FA",
"TEMPORARY_HOLD",
"TERMINATE_SESSION",
],
"description": "Immediate mitigation action",
},
}
def main():
schema = StructuredSchema(schema_definition)
context = """
User ID: usr_9921
Location: Lagos, Nigeria (usual: Seattle, USA)
Device: Unknown Linux Chromium browser
Action: Wire transfer $49,500 to offshore escrow
Prior velocity: 0 transfers in 90 days
"""
result = run_parallel_generation(context, schema)
print(f"Elapsed time: {result['elapsed_ms']} ms")
print(f"Sequential passes: {result['sequential_forward_passes']}")
print(f"Parsed JSON: {result['parsed_json']}")
if __name__ == "__main__":
main()Example output:
Elapsed time: 146.88 ms
Sequential passes: 1
Parsed JSON: {
'fraud_risk': {'value': 'SUSPICIOUS', 'prob': 0.8131},
'block_account': {'value': True, 'prob': 0.6637},
'recommended_action': {
'value': 'TEMPORARY_HOLD',
'prob': 0.6054
}
}
Set MODEL_ID before running your application:
export MODEL_ID=mlx-community/Qwen2.5-0.5B-Instruct-4bitFor example:
MODEL_ID=mlx-community/Qwen2.5-0.5B-Instruct-4bit python example.pyIf MODEL_ID is not set, the engine uses:
mlx-community/Qwen2.5-1.5B-Instruct
The selected model must be compatible with the active inference backend.
Schemas are created with StructuredSchema. Each field includes:
type: eitherenumorbooleandescription: semantic guidance for the modelchoices: allowed values for enum fields
{
"priority": {
"type": "enum",
"choices": ["LOW", "MEDIUM", "HIGH", "CRITICAL"],
"description": "Urgency of the incident",
}
}{
"requires_review": {
"type": "boolean",
"description": "Whether the result requires manual review",
}
}from pcd_rlcd.schema import StructuredSchema
schema = StructuredSchema(
{
"priority": {
"type": "enum",
"choices": ["LOW", "MEDIUM", "HIGH", "CRITICAL"],
"description": "Urgency of the incident",
},
"requires_review": {
"type": "boolean",
"description": "Whether manual review is required",
},
"department": {
"type": "enum",
"choices": [
"BILLING",
"INFRASTRUCTURE",
"SECURITY",
"PRODUCT_SUPPORT",
],
"description": "Department responsible for handling the issue",
},
}
)run_parallel_generation returns a dictionary containing the structured result and execution telemetry.
{
"mode": "parallel_constrained_calibrated",
"elapsed_ms": 146.88,
"prefill_ms": 100.42,
"suffix_eval_ms": 42.17,
"sequential_forward_passes": 1,
"is_valid_json": True,
"schema_match": True,
"parsed_json": {
"priority": {
"value": "HIGH",
"prob": 0.91,
},
"requires_review": {
"value": True,
"prob": 0.84,
},
"department": {
"value": "SECURITY",
"prob": 0.96,
},
},
"field_telemetry": {
"priority": {
"value": "HIGH",
"confidence": 0.91,
"cardinality": 4,
"top_choices": [
{
"choice": "HIGH",
"probability": 0.91,
},
{
"choice": "CRITICAL",
"probability": 0.06,
},
],
}
},
}The exact timing fields depend on the selected backend, model, hardware, and input.
Traditional structured generation produces JSON autoregressively:
"context" -> "{" -> "\"field\"" -> ":" -> "\"value\"" -> ...
Each generated token requires another model evaluation. As the output grows, the number of sequential evaluations also grows.
Parallel constrained decoding uses a different strategy:
- The input context and schema descriptions are prefetched once.
- The resulting key-value cache is shared across schema fields.
- Each field is evaluated against only its valid candidate values.
- Candidate probabilities are normalized within each field.
- The selected values are assembled programmatically.
- The final structure is guaranteed to match the requested schema.
This approach is especially useful when fields contain bounded values such as enums, labels, routing decisions, or booleans.
.
├── pyproject.toml
├── README.md
├── requirements.txt
├── uv.lock
└── src
└── pcd_rlcd
├── __init__.py
├── benchmark.py
├── engine.py
├── engine_mlx.py
├── engine_torch.py
├── prompt_builder.py
├── py.typed
└── schema.py
schema.py— schema definitions and field metadataengine.py— public engine interfaceengine_mlx.py— MLX-based inference implementationengine_torch.py— PyTorch-based inference implementationprompt_builder.py— prompt construction utilitiesbenchmark.py— benchmark and comparison utilities
The package includes benchmark utilities for comparing parallel constrained decoding with autoregressive generation.
Run the benchmark module with:
python -m pcd_rlcd.benchmarkBenchmark results depend on:
- Model size and quantization
- Apple Silicon generation
- Available memory
- Input length
- Number of fields
- Number of candidate choices
- Backend configuration
Use benchmark results as hardware- and model-specific measurements rather than universal performance guarantees.
Enum fields work best when:
- The candidate set is known in advance
- Each choice is semantically distinct
- Values are reasonably short
- The model can infer the correct choice from the context
For example:
{
"severity": {
"type": "enum",
"choices": ["INFO", "WARNING", "ERROR", "CRITICAL"],
"description": "Operational severity of the event",
}
}For open-ended text generation, use a conventional autoregressive generation approach instead.
- The engine is designed for bounded structured decisions, not arbitrary JSON generation.
- Results are model predictions and should be validated against application-specific requirements.
- Confidence values represent model probabilities over the candidate set; they are not guaranteed to be calibrated for every domain.
- Performance varies significantly between models, devices, and backends.
- MLX execution requires compatible Apple Silicon hardware.
- The available schema field types are currently limited to the types supported by the package implementation.
This package extracts and refactors the reusable core of the original Parallel Constrained Decoding implementation into an independently installable Python package.
The original project demonstrated parallel constrained decoding for Apple Silicon using MLX. pcd-rlcd focuses on making the core schema and inference functionality easier to install and use from other Python applications.
Apache License 2.0.