A100 fixed-demand POC complete · 3 of 3 pairs passed

AI inference capacity control

More useful AI service. Same GPU infrastructure.

S2 is a software control platform between AI applications and model serving. It is designed to reduce avoidable physical model calls—fresh execution requests sent to the model server—so more accepted demand can be served from available GPU capacity.

For infrastructure teams facing retries, repeated demand, obsolete agent work, queue pressure or a near-term GPU expansion decision.

Outcome-focused. Quality-gated. Fail-open. Evidence-led.

Application requests passing through a governed S2 layer so only necessary generation reaches the GPU and accepted results return.
Logical demandS2 capacity controlRequired GPU work
Same model obligationRequired work forwardedQuality-gated outcomesGPU-board boundaryEvidence retained

Completed evidence

A bounded A100 matched-pair result, presented with its limits.

Review the full evidence position
26.19% fewerPhysical vLLM calls

Observed across three valid fixed-demand pairs at approximately 24.99 logical requests per second.

1.76% lowerGPU energy per accepted request

Measured at the GPU-board boundary across the same three A100 pairs.

3 of 3Matched pairs passed

Every pair retained effectively equivalent accepted service with stable queues.

Self-run Google Colab POC on a full A100-SXM4-80GB using Qwen2.5-7B-Instruct BF16 and vLLM 0.23.0. Indicative paired evidence only; no measured capacity uplift or production claim.

01

What S2 is

Capacity control for the demand before your model server.

Applications ask for outcomes. Model servers perform physical work. S2 sits between them to reduce avoidable calls while preserving the same model obligation—the agreed model and service requirement that must still be fulfilled.

A practical example

Requested service and necessary GPU work are not always the same.

Several application requests may compete for GPU execution even though only one accepted outcome is ultimately useful. S2 controls eligible demand before unnecessary physical work reaches the GPU, while forwarding work that still requires model execution.

Logical request: a service obligation presented by an application. Physical model call: a fresh execution request sent to the model-serving environment.

AI applications sending requests through S2 before required work reaches a model server and GPU.
S2 is evaluated at the boundary between application demand and model-server compute.

The problem

Logical demand is not always necessary physical work.

Retries, repeated demand, simultaneous requests and obsolete agent work can add queue pressure without adding equal business value.

The platform

S2 controls avoidable demand before GPU execution.

Necessary work continues to the model. Accepted outcomes remain subject to the agreed quality and service requirements.

The decision

Test whether software can recover headroom first.

An evidence-led assessment can inform capacity planning before additional accelerators, power and cooling are committed.

02

Choose your route

Move directly to the decision you need to make.

The homepage establishes the position. The dedicated pages carry the technical, evidence and engagement detail.

01

Understand the platform

See where S2 sits, what work remains protected and which demand patterns may contain recoverable capacity.

Explore Platform
02

Review the evidence

Examine the completed A100 fixed-demand POC, laptop result and explicit limits attached to each evidence stratum.

Review Evidence
03

Assess your workload

Prepare a non-confidential request and determine whether a controlled technical assessment is justified.

Start an Assessment
03

Workload fit

Is S2 relevant to your AI workload?

S2 may be relevant when recoverable demand contributes to capacity, latency or queue pressure. The question is assessed against the workload—not assumed from a generic benchmark.

Frequent retriesRepeated or substantially equivalent requestsSimultaneous demand for the same outcomeAgent work that becomes obsoleteQueue pressure during spikesGPU expansion under consideration

Honest limitation Workloads composed almost entirely of unique, necessary requests may contain less recoverable capacity.

Company

Built by ProgGen Zeroth AI Limited.

ProgGen develops evidence-led systems intended to improve useful service from computing infrastructure. Demonstrated results remain separate from intended benefits, and every claim stays attached to its measurement boundary.

Company and operating principles
Legal name
PROGGEN ZEROTH AI LIMITED
Company number
17374280
Company type
Private company limited by shares
Jurisdiction
England and Wales
Incorporated
31 July 2026

The next step

Assess the workload. Prove the boundary.

Start with non-sensitive workload characteristics. Move to technical detail only inside the right confidential engagement.

Designed for

  • Infrastructure leaders
  • Inference platform owners
  • Capacity-constrained AI teams