Build agents on the
OreoFlow runtime.
A small, typed Python boundary for model routing, OpenAI-compatible calls, streaming, tool schemas, budgets, and usage accounting.
A stable runtime boundary
Applications own agent behavior. OreoFlow owns the model-facing mechanics that every agent otherwise repeats.
Compatibility ruleConsumer applications import from oreoflow, not internal rtk modules. Internal code can evolve without breaking application imports.
Pin the released framework
Use the Git tag in production and an editable sibling checkout while developing both repositories.
Released tag
python -m pip install \
"git+https://github.com/elixpo/agent.elixpo.git@v0.2.0"Local development
python -m pip install -e ../agent.elixpoThe distribution name is elixpoo. The supported application import is oreoflow. Search's Docker build installs the sibling repository through a named BuildKit context.
Route capabilities, not model names
Agents request a logical role. A small YAML file decides which provider model serves it.
base_url: https://gen.pollinations.ai/v1
defaults:
effort: low
roles:
classify: {model: nova-fast}
code: {model: qwen-coder}
prose: {model: nova-fast}Keep the keys separate
- POLLINATIONS_API_KEY
- Application → provider
- API_KEY
- User → your public API
OreoFlow receives provider credentials explicitly. It does not discover or retain an application's environment file.
The v0.2.0 public surface
These names are re-exported from oreoflow and form the current compatibility contract.
RouterResolves logical roles to models and performs calls or streams.
MessageValidated OpenAI-compatible system, user, assistant, and tool messages.
ToolDefOpenAI function-tool declaration. Execution remains application-owned.
BudgetIn-memory per-task token budget with an absolute runaway ceiling.
TokenLedgerOptional append-only JSONL usage records for completed calls.
LLMClientLow-level asynchronous OpenAI-compatible HTTP client with retries.
UsagePrompt, cached, completion, and total token accounting.
EffortLow, medium, or high effort mapped to controlled temperatures.
One router, calls or chunks
A router keeps one HTTP client per selected model. Reuse it for related work and close it during shutdown.
import asyncio, os
from dotenv import load_dotenv
from oreoflow import Budget, Message, Router, load_models_config
async def main():
load_dotenv(".env.local")
router = Router(
"task-42",
models=load_models_config("models.yaml"),
api_key=os.environ["POLLINATIONS_API_KEY"],
budget=Budget("task-42", limit=4_000),
)
try:
response = await router.call(
"code",
[Message(role="user", content="Write a URL validator")],
effort="low",
max_tokens=500,
)
print(response.choices[0].message.content)
finally:
await router.aclose()
asyncio.run(main())Stream provider chunks
async for chunk in router.stream("prose", messages, effort="low"):
for choice in chunk.choices:
if choice.delta.content:
print(choice.delta.content, end="", flush=True)The framework yields validated chunks immediately. Your HTTP application decides whether to forward every delta, batch characters, or convert output to SSE events.
Declare in OreoFlow, execute in your app
The runtime validates OpenAI function schemas. It intentionally does not authorize or execute application functions.
tool = ToolDef.model_validate({
"type": "function",
"function": {
"name": "lookup_weather",
"description": "Read weather for one city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
})Bound every model task
The soft limit is visible to the application; the multiplied hard ceiling is the runaway kill switch.
remaining() reports capacity. Crossing the soft limit is advisory.
A pre-call estimate past limit × kill_multiple raises BudgetExceeded.
Completed calls append model, role, cached, prompt, completion, and total usage to JSONL.
budget = Budget("task-42", limit=10_000, kill_multiple=3)
ledger = TokenLedger("state/token_log.jsonl")
router = Router("task-42", models=models, api_key=key,
budget=budget, ledger=ledger)How Search uses OreoFlow
The deployed Search agent crosses this boundary for every live specialist call and stream.
Search imports Message, Router, and ToolDef from oreoflow. It owns deterministic agent selection, skill injection, Redis response chains, Qdrant memory, and transport buffering.
User-selectable response effort
Search maps the Responses API reasoning control directly into OreoFlow for decision and specialist calls.
curl https://search.elixpo.com/v1/responses \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "writing",
"input": "Write a concise release announcement.",
"reasoning": {"effort": "medium"},
"stream": true
}'What v0.2.0 does not provide
These remain application responsibilities or roadmap work—not hidden framework features.
- Multi-agent scheduling, rooms, floors, delegation, or A2A transport
- A skill registry, policy engine, or tool executor
- Redis/Qdrant memory and OpenAI conversation persistence
- Image, PDF, browser, or web-search implementations
- HTTP endpoints, SSE batching, or public client authentication
- A shared router daemon across independent processes
Raw Python tests
Resolve roles without network access, then opt into one bounded live provider call.
Framework
cd ~/agent.elixpo
python examples/raw_oreoflow.py --role code
python examples/raw_oreoflow.py \
--role prose --live --stream \
--env-file ../search.elixpo/.env.local \
--prompt "Reply with exactly: ready."Search agent
cd ~/search.elixpo
python tester/raw_agent_runtime.py \
--agent coding \
--effort medium \
--live --stream \
--prompt "Write a URL validator"