Introducing Endeavor 1.0 from Flower Labs



Frontier AI, under your control
For organizations adopting frontier AI, capability has usually come with dependency. The most powerful models are typically accessed through closed APIs, which makes them convenient to use but leaves the products, agents, and workflows built around them dependent on a single provider. Models that can be deployed independently offer greater control, but until now that has often required a meaningful compromise in capability.
Today, we are introducing Endeavor 1.0, our most powerful model yet. Endeavor is a frontier-class generalist with strong reasoning and coding performance, designed for the long-horizon work increasingly carried out by agents. It is competitive with leading models from OpenAI and Anthropic, while remaining available as a service operated by Flower or for deployment inside infrastructure an organization controls.
Endeavor 1.0 is our first production-ready release of Endeavor. We are launching it initially as a preview, onboarding a select group of organizations and partners rather than opening immediately for full self-service model access. This will allow us to work closely with early adopting organizations as we also expand compute availability.
Our goal with Endeavor is not simply to provide another model endpoint. We want organizations to be able to start with frontier capability and build a lasting AI foundation around it, extending the model with their own agents, evaluations, data pipelines, and improvement loops over time.
From Lizzy to the frontier
Endeavor 1.0 arrives just four months after Lizzy, Flower's sovereign 7B model. Lizzy was built for UK use and demonstrated how a model can combine specialized knowledge and reasoning with particular deployment requirements. Endeavor is the next step in that program. It is a broad generalist designed to compete directly with the leading frontier systems across reasoning, software development, and complex knowledge work.
Endeavor's initial evaluation results place it firmly in that group. It achieves 92.0 on GPQA, 98.2 on HumanEval, 99.9 on AIME 2026, and 94.1 on IFEval.
| GPQA | HumanEval | AIME 2026 | IFEval |
|---|---|---|---|
| 92.0 | 98.2 | 99.9 | 94.1 |
Across the four completed core evaluations included in our launch comparison, Endeavor records the highest HumanEval score. It matches GPT-5.6 Sol and Claude Fable 5 on AIME 2026, outperforms Kimi K3 on three of the four benchmarks, and outperforms Nemotron 3 Ultra on all four.
| Benchmark | Endeavor 1.0 | GPT-5.6 Sol | Claude Fable 5 | Kimi K3 | Nemotron 3 Ultra |
|---|---|---|---|---|---|
| GPQA | 92.0 | 94.1 | 92.6 | 93.5 | 86.7 |
| HumanEval | 98.2 | 95.1 | 97.0 | 96.3 | 96.3 |
| IFEval | 94.1 | 95.9 | 91.7 | 92.8 | 91.9 |
| AIME 2026 | 99.9 | 99.9 | 99.9 | 96.7 | 94.2 |
No small collection of benchmarks can fully describe how useful a model will be in practice. These results do, however, establish Endeavor as a serious new entrant at the frontier and demonstrate the breadth we have aimed for. Endeavor is not narrowly optimized for a single benchmark family. It is intended to serve as the core generalist model across products, agents, and complex workflows.
A foundation, not just an endpoint
For many organizations, a managed model service is the fastest path to production. Flower will operate Endeavor as a production service for customers that want us to manage deployment, scaling, and model operations. Choosing a managed endpoint, however, should not have to become a permanent infrastructure decision. Endeavor can also be deployed inside an organization's own environment, with Flower supporting its integration across existing data, compute, and application systems.
Organizations can use Flower's service for most workloads while deploying selected sensitive applications privately. They can also move a broader deployment into their own infrastructure as their needs evolve. This gives teams the freedom to choose the operating model that fits their immediate requirements without making that first choice permanent.
That freedom matters more as the work around a model accumulates. Building a serious AI capability involves far more than writing prompts. Over time, organizations develop agents, evaluations, tool integrations, data pipelines, guardrails, and feedback systems around the underlying model. With Endeavor, those investments do not have to remain permanently tied to a single closed API. Teams can maintain a stable deployment, introduce improvements on their own timetable, and preserve a path to operate the capability independently of any one provider.
We believe frontier AI should be something organizations can build on, not only something they rent.
Built from a different starting point
Flower has spent years building the model, training, and distributed systems technology required for powerful AI that organizations can trust, deploy on infrastructure they choose, and improve around their own work. Endeavor brings that foundation to the frontier.
A modern frontier model is more than its weights. Practical performance depends on the complete system around the model, and a major part of our work on Endeavor focuses on how the model interacts with the harness around it. This includes how it allocates reasoning effort, constructs and preserves context, interprets tool interfaces, uses tools, and checks, revises, and recovers when a step fails.
We also bring in signals that conventional public benchmarks cannot provide. One of those comes from FlowerBench, our benchmark for long-horizon enterprise work. FlowerBench draws only on tasks contributed by organizations that have explicitly opted in to participate through the Flower Enterprise Evaluation Network. The tasks run inside those organizations' own environments, so proprietary data and internal context remain in place, while sanitized results reveal how models perform on work shaped by real tools, domain rules, and deliverables. These signals have helped guide Endeavor's evaluation design, harness development, post-training, and wider system improvements.
Our model development deliberately builds on the progress of the broader open-weight ecosystem. Endeavor draws on mature, widely available capabilities that leading open-weight models already provide well, including general language understanding, broad public knowledge, and common coding patterns. Rather than recreating those foundations, we build on them with capabilities developed through Flower's own model program. These include UK-specialist knowledge and reasoning developed through Lizzy, as well as new model behaviors, specialist capabilities, and training advances developed specifically for Endeavor. Continual pre-training, targeted post-training, and careful model integration bring these strengths together and will enable us to keep refining Endeavor over time.
Using Endeavor 1.0
Endeavor can be called from existing applications and agent frameworks through familiar model APIs and response formats. For most organizations, the simplest route will be Flower's managed service, where we handle deployment, scaling, and model operations. Private deployment is also available for organizations that need Endeavor to run inside their own environment or want greater control over sensitive workloads and infrastructure.
During the Endeavor 1.0 preview, access is available by request. We are onboarding a limited number of organizations and partners for managed and private deployments, and will broaden availability over time.
Endeavor offers a different path from the usual choice between a high-capability closed API and a more controllable but less capable model. It begins with frontier performance while giving organizations a credible route to make that intelligence a durable part of their own technology.
Request access to Endeavor 1.0, extend it around your own work, and build intelligence that becomes increasingly your own.