@tauska67/agentapp-builder
flwr new @tauska67/agentapp-builderAgentApp Builder
Collaborative Endeavor agents generate independent Flower AgentApps.
A Planner freezes the acceptance plan, a Developer generates an application specification, and a QA Tester reviews the frozen candidate against executed checks. Rejected candidates receive a bounded repair attempt. An accepted candidate exports as a standalone Python project with its own Flower entry point, runtime, schemas, tests and license.
All real model requests use flower-endeavor. The project follows the runtime interfaces in Flower's @flwrlabs/hackathon-collab-agent-recipe and its linked tutorial. The local fixture mode is explicit, deterministic, and makes zero model requests.
Delivery status
Implemented: generation orchestration, role-specific contexts, acyclic child execution, protected contract tests, repair feedback, source hashing, independent export, local web dashboard, CLI-to-SuperGrid bridge, chunked result import, and hash-bound publication approval. Local verification details are in verification/REPORT.md.
Verified on the event machine on 16 September 2026 (see verification/REPORT.md): dependency resolution with a retained uv.lock, both FAB builds, the offline fixture build in the dashboard, the exported child in a separate process, and the SuperGrid smoke check with Endeavor (flower-endeavor resolves to flower-endeavor-v1.0): text reply, one real web_search connector call, argument validation, and the model's receipt of the tool result. A full live generator build with a live child preview was executed on Google Gemini through the direct transport (verification/gemini-build.json): one candidate accepted, 6 model calls, 52 seconds. The same build on SuperGrid with Endeavor and web connectors has not been executed yet because it costs credits; see "Costs" below. No app or repository has been published by this delivery.
Costs
SuperGrid charges credits per token. The smoke check (three Endeavor calls, 7.4k input and 1.6k output tokens, one connector call) cost 23 credits; the first attempt cost 30. Endeavor answered in 11 to 15 seconds per call with the default reasoning.effort=low. A full build with a live child preview makes roughly 8 to 12 calls and is estimated at 200 to 350 credits; builder.live-check=false reduces it to 3 calls per candidate. The per-build caps (builder.max-model-calls, builder.max-seconds) are the cost guard.
1. Start the offline demonstration
Requirements: Python 3.11 through 3.13 and uv. Dependency installation requires network access. Run from this repository directory.
./scripts/start.sh demo
Open http://127.0.0.1:8765 and select Generate AgentApp. The disclosed fixture intentionally produces a first candidate with a missing source-failure warning. Actual runtime contract execution rejects it; the second fixture candidate repairs it. Inspect Activity, Validation, Candidate diff and Source files, then run the exported child in a separate Python process or export its .tar.gz.
The fixture uses one fixed demonstration requirement. It does not generate arbitrary applications and provides no evidence of Endeavor's model quality.
For a command-line demonstration:
uv sync --extra ui --extra dev uv run agentapp-builder demo --out .runs/first-demo cd .runs/first-demo/technical-source-comparison uv run python -m generated_app --self-test uv run python -m generated_app --mode fixture
Existing output directories are never overwritten. Choose a new output directory for each run. With dependencies already installed, the same commands work as python -m agentapp_builder ... from the repository root.
2. Start live generation with Endeavor
./scripts/start.sh live
This installs dependencies, starts Flower's interactive login, performs the text and one-tool smoke check, and starts the local dashboard. Choose Live / SuperGrid + Endeavor. The dashboard submits the generator to SuperGrid through your authenticated Flower CLI. It receives role artifacts and the final hash-checked result from the run log.
The manual equivalent is:
uv sync --extra ui --extra dev uv run flwr build uv run flwr login supergrid uv run flwr run . supergrid --run-config 'agent.action="smoke"' --stream uv run agentapp-builder ui
A failed smoke test stops the startup script. Check whether the event federation grants flower-endeavor, the Responses endpoint, web_search and web_fetch. There is no model fallback. The SDK uses max_retries=0; failed requests are not retried automatically.
Live CLI generation:
uv run agentapp-builder build \ --requirement 'Create a three-agent comparison app with research, verification and reporting. Use public web tools. Report unavailable sources. Return summary and findings with source_refs.' \ --transport supergrid \ --out .runs/live-001
The default build limits are 3 candidates, 18 model requests, 12 connector calls, and 900 seconds of cooperative elapsed-time checks. Each model request has a timeout of at most 240 seconds (Endeavor answered builder-sized prompts in 10 to 90 seconds). Runtime tool timeouts are owned by the Flower connector; a blocked connector can exceed the cooperative elapsed-time limit. Dependency installation and remote queue time are additional. No completion time or model-quality result is assumed.
2b. Real-world generation on another provider (no SuperGrid credits)
The direct transport can target any OpenAI-compatible provider. Chat-only providers such as Google Gemini are wrapped by agentapp_builder/chat_adapter.py (also copied into every generated app). Web connectors exist only on SuperGrid, so use --no-web; the child preview then runs with Gemini too.
export AAB_MODEL_API=chat AAB_MODEL_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai/ export AAB_MODEL=gemini-3.8-flash AAB_MODEL_API_KEY=... # never commit the key uv run agentapp-builder build --requirement "$(sed -n '/## Meeting note/,/^For this/p' examples/requirements.md | sed '1d;$d')" \ --transport direct --no-web --out .runs/gemini-001
On SuperGrid the model is always flower-endeavor; the requested model is recorded in every call record.
Zero-credit local run of the SuperGrid path
A local Flower SuperLink can execute this AgentApp with Google Gemini as the model provider, so the exact code that runs on SuperGrid can be rehearsed without credits. The gemini-responses-proxy folder next to this project exposes an Open Responses endpoint in front of Gemini:
(cd ../gemini-responses-proxy && GEMINI_API_KEY=... ./run_local_stack.sh) # proxy on 9100, SuperLink on 9091 uv run flwr run . local-agent --run-config 'agent.input="Create a two-agent app that summarises pasted text" builder.allow-web=false' --stream
Catalog model IDs such as flower-endeavor-v1.0 are mapped to Gemini by the proxy. Flower's built-in web connectors are not available on the local runtime.
3. Inspect and independently run a generated application
cd .runs/live-001/technical-source-comparison # Use the actual exported name. uv sync --extra dev uv run python -m generated_app --self-test uv run pytest -q uv run flwr build uv run flwr run . supergrid --stream
A generated child imports only its own generated_app package and declared dependencies. Its standalone live run receives new user input and separate Flower runtime credentials. The parent's live preview and this independent deployment are separate verification steps.
The local dashboard's Run exported child action runs the exported package in a separate Python process for fixtures, or submits that exported project as a separate SuperGrid run for live mode. The dashboard itself remains local; Flower Hub exposes the AgentApp chat and structured outputs. The dashboard is not hosted on Hub by this implementation.
4. Run the generator directly on SuperGrid and recover its result
set -o pipefail uv run flwr run . supergrid --stream 2>&1 | tee generator.log uv run agentapp-builder import-log generator.log --out .runs/imported-001
AAB_RESULT_V1 log records carry numbered base64 chunks with integrity hashes. Missing, conflicting or modified chunks fail import. Hashes establish integrity relative to the received result; they do not authenticate a malicious third-party log producer. Do not import logs from untrusted users.
For a refinement in the same Flower run series, set builder.resume=true. The previous accepted AppSpec is loaded from Context; the new requirement creates a new acceptance plan. Criteria remain frozen within each build. There is no model training phase.
5. Publish reviewed source
Publishing is a separate CLI action. Review the exact staged source, including requirement-derived instructions and sample inputs, before approval.
uv run agentapp-builder stage \ --project . \ --publisher YOUR_FLOWER_USERNAME \ --out dist/publish-builder
This prints a source_hash and creates a sibling approval manifest. Replace APPROVED_SHA256 with that full hash after reviewing every staged file:
# Verification only. No network write. uv run agentapp-builder publish \ --stage dist/publish-builder \ --approved-digest APPROVED_SHA256 # Explicit public-source publication of the approved bytes. uv run agentapp-builder publish \ --stage dist/publish-builder \ --approved-digest APPROVED_SHA256 \ --execute
Changed, added or symlinked staged files invalidate approval. Publication builds a separate copy and publishes a clean copy of the approved source. Flower's underlying publish command uploads immediately, so the wrapper performs review and integrity checks before invoking it. The wrapper never creates a GitHub repository or pushes commits. Publish this repository to your chosen Git hosting service yourself after inspecting the public files.
For a child app, pass its exported directory to stage --project. Set the actual publisher; local is a placeholder and cannot pass staging.
Dashboard on another port
./scripts/serve_dashboard.sh demo 8871 # offline fixtures ./scripts/serve_dashboard.sh live 8871 # SuperGrid + Endeavor through your Flower CLI login
Verification and development
./scripts/verify.sh uv run agentapp-builder doctor uv run ruff check .
verify.sh runs the tests, exports a child, runs its self-tests from that directory and attempts FAB builds for both projects. It makes no model request and publishes nothing. uv.lock was produced by a real uv sync on 16 September 2026 (Flower 1.37.0, OpenAI SDK 2.x); keep it so the event environment resolves the same dependency set.
Supported generation scope
Text input, structured JSON output, 2 to 6 agents, an acyclic dependency graph, typed outputs, and optional read-only web_search / web_fetch. The model generates role instructions, schemas, dependencies and tool selection. A deterministic renderer writes reviewable Python code from fixed runtime modules. Candidate-provided Python, arbitrary shell execution, arbitrary dependencies, external account writes, purchases, recursive publishing and model training are outside this version's scope. Unsupported requirements and essential clarification questions stop generation explicitly.
Files
agentapp_builder/ agent_app.py Flower generator entry point and smoke action engine.py Planner / Developer / QA state transitions spec.py AppSpec, acceptance plan, schemas and limits child.py Standalone child DAG executor model.py Endeavor Responses adapter and accounting validation.py Protected executable contract tests render.py Deterministic standalone project export publish.py Reviewed staging and exact-digest publication protocol.py Hash-checked result transfer through logs supergrid.py Authenticated Flower CLI bridge server.py Loopback dashboard API static/ Local HTML, CSS and JavaScript fixtures.py Explicit deterministic test doubles examples/ Requirements, reference spec and recorded fixture evidence docs/ Architecture, security, event runbook and primary sources tests/ Local tests and documented-interface test doubles verification/ Tests actually executed for this archive
Read docs/ARCHITECTURE.md, docs/SECURITY.md, docs/EVENT_RUNBOOK.md and docs/SOURCES.md for implementation details and provenance.