Observed across three valid fixed-demand pairs at approximately 24.99 logical requests per second.
AI inference capacity control
More useful AI service. Same GPU infrastructure.
S2 is a software control platform between AI applications and model serving. It is designed to reduce avoidable physical model calls—fresh execution requests sent to the model server—so more accepted demand can be served from available GPU capacity.
For infrastructure teams facing retries, repeated demand, obsolete agent work, queue pressure or a near-term GPU expansion decision.
Outcome-focused. Quality-gated. Fail-open. Evidence-led.

Completed evidence
A bounded A100 matched-pair result, presented with its limits.
Review the full evidence positionMeasured at the GPU-board boundary across the same three A100 pairs.
Every pair retained effectively equivalent accepted service with stable queues.
Self-run Google Colab POC on a full A100-SXM4-80GB using Qwen2.5-7B-Instruct BF16 and vLLM 0.23.0. Indicative paired evidence only; no measured capacity uplift or production claim.
What S2 is
Capacity control for the demand before your model server.
Applications ask for outcomes. Model servers perform physical work. S2 sits between them to reduce avoidable calls while preserving the same model obligation—the agreed model and service requirement that must still be fulfilled.
A practical example
Requested service and necessary GPU work are not always the same.
Several application requests may compete for GPU execution even though only one accepted outcome is ultimately useful. S2 controls eligible demand before unnecessary physical work reaches the GPU, while forwarding work that still requires model execution.
Logical request: a service obligation presented by an application. Physical model call: a fresh execution request sent to the model-serving environment.

The problem
Logical demand is not always necessary physical work.
Retries, repeated demand, simultaneous requests and obsolete agent work can add queue pressure without adding equal business value.
The platform
S2 controls avoidable demand before GPU execution.
Necessary work continues to the model. Accepted outcomes remain subject to the agreed quality and service requirements.
The decision
Test whether software can recover headroom first.
An evidence-led assessment can inform capacity planning before additional accelerators, power and cooling are committed.
Choose your route
Move directly to the decision you need to make.
The homepage establishes the position. The dedicated pages carry the technical, evidence and engagement detail.
Understand the platform
See where S2 sits, what work remains protected and which demand patterns may contain recoverable capacity.
Explore PlatformReview the evidence
Examine the completed A100 fixed-demand POC, laptop result and explicit limits attached to each evidence stratum.
Review EvidenceAssess your workload
Prepare a non-confidential request and determine whether a controlled technical assessment is justified.
Start an AssessmentWorkload fit
Is S2 relevant to your AI workload?
S2 may be relevant when recoverable demand contributes to capacity, latency or queue pressure. The question is assessed against the workload—not assumed from a generic benchmark.
Honest limitation Workloads composed almost entirely of unique, necessary requests may contain less recoverable capacity.
Company
Built by ProgGen Zeroth AI Limited.
ProgGen develops evidence-led systems intended to improve useful service from computing infrastructure. Demonstrated results remain separate from intended benefits, and every claim stays attached to its measurement boundary.
Company and operating principles- Legal name
- PROGGEN ZEROTH AI LIMITED
- Company number
- 17374280
- Company type
- Private company limited by shares
- Jurisdiction
- England and Wales
- Incorporated
- 31 July 2026
The next step
Assess the workload. Prove the boundary.
Start with non-sensitive workload characteristics. Move to technical detail only inside the right confidential engagement.
Designed for
- Infrastructure leaders
- Inference platform owners
- Capacity-constrained AI teams