What is S2?+
S2 is a software capacity-control platform placed between AI applications and model serving. It is designed to reduce avoidable physical model calls while preserving accepted service outcomes.
Who is S2 for?+
S2 is for organisations operating AI inference workloads with observable capacity, latency or queue pressure.
What problem does it solve?+
It addresses the gap between logical demand and necessary physical computation. Retries, repeated demand, simultaneous demand and obsolete work can consume GPU capacity without creating equal value.
What is a logical request?+
A logical request is a service obligation presented by an application. It counts as useful service only when the delivered outcome meets the agreed requirements.
What is a physical model call?+
A physical model call is a fresh execution request sent to the model-serving environment. S2 measures how many fresh calls are required to serve accepted logical demand.
What is quality-accepted service?+
It is a delivered outcome that passes the quality and service requirements agreed for the workload. Rejected outcomes do not count as useful capacity.
Does S2 replace the existing model server?+
No. S2 is positioned before the existing model-serving environment. Required work continues to the agreed model and serving platform.
Does S2 require model retraining?+
The public S2 position does not require changing or retraining the model. The model obligation remains unchanged.
Where does S2 sit in the serving environment?+
S2 sits at the boundary between application demand and model serving. It is evaluated as a controlled service adjacent to the existing inference platform.
What remains under the customer’s control?+
The customer retains control of model choice, serving environment, workload and tenant boundaries, quality requirements, service objectives and deployment approval.
Does S2 change the model obligation?+
No. The model obligation is the agreed model and service requirement that must still be fulfilled. Work that requires fresh model execution is forwarded.
Does it make the GPU faster?+
No. S2 is intended to increase useful service from available capacity by reducing avoidable work, not by changing GPU performance.
What happens when optimisation is unavailable?+
S2 fails open: required work is forwarded when an optimisation cannot be used safely or is unavailable.
Can S2 help a workload made entirely of unique requests?+
Opportunity may be limited when almost every request is unique and necessary. The assessment may conclude that S2 is not the right next step.
How is an acceptable outcome defined?+
The customer and assessment team agree workload-specific quality and service requirements before comparison. A result is useful only when it passes those requirements.
What does the A100 fixed-demand POC show?+
In one self-run Google Colab session using a full A100-SXM4-80GB, Qwen2.5-7B-Instruct BF16 and vLLM 0.23.0, three valid matched pairs retained effectively equivalent quality-accepted service while Full S2 used 26.19% fewer physical model calls and approximately 1.76% less GPU-board energy per accepted request.
Does the A100 result prove a 35.48% capacity uplift?+
No. The 35.48% figure is a calculated physical-work-equivalent indicator. The POC did not search for native or S2 capacity ceilings, admit additional logical demand or establish a measured capacity uplift.
What are the principal A100 limitations?+
The result comes from one self-run Colab session and three deliberately controlled fixed-demand pairs. The protocol did not include measured live unique-work negative controls or sensitivity conditions, and its 20% retry and 10% fan-out traffic profile is not assumed to represent an individual customer.
What does the completed laptop evidence show?+
Across five valid fixed-demand pairs covering four qualified Ollama models, S2 used 28.19% fewer physical calls and 16.10% less GPU-board energy per accepted request while serving equivalent accepted demand.
What capacity result has been established?+
For Llama 3.1 8B Q4_K_M through Ollama 0.31.1 on the tested RTX 5050 laptop, S2 sustained at least 50.09% more quality-accepted logical workload than the confirmed native ceiling.
What energy boundary is measured?+
The completed validation measures GPU-board energy at the GPU device boundary. It does not claim whole-laptop, whole-server, rack, cooling or facility energy savings.
Why did latency not improve in every comparison?+
S2 is not presented as a universal latency reduction. The valid comparisons maintained quality and service attainment, but latency did not improve in every pair.
Which serving environments have been demonstrated?+
The completed evidence contains separate strata: Ollama on an RTX 5050 laptop across four qualified models, and a fixed-demand POC using Qwen2.5-7B-Instruct BF16 through vLLM 0.23.0 on a full A100-SXM4-80GB hosted by Google Colab. The two strata are not pooled.
Is S2 production-ready?+
The site does not claim production readiness. Production-cluster validation, independent reproduction and customer-specific operational assurance have not been completed.
Is the current validation independent?+
No. Both the laptop programme and the hosted A100 POC are self-run. Independent reproduction and production-cluster validation have not been completed.
Has A100 or vLLM testing been completed?+
Yes, within a narrow POC boundary. Three valid fixed-demand pairs were completed through vLLM 0.23.0 on a full Google Colab A100 allocation. This is indicative paired evidence, not a measured A100 capacity ceiling or production validation.
How can an organisation request an assessment?+
Prepare a non-confidential request on the assessment page. The website opens a draft addressed to info@proggen.co.uk; nothing is sent until the visitor reviews and sends it through their email service.
What happens after an assessment request is sent?+
ProgGen reviews the non-confidential information, discusses workload fit and, if further evaluation is justified, moves technical detail into an appropriate confidential engagement. A pilot is not guaranteed.
What information should not be sent by email?+
Do not send confidential workload data, source code, credentials, raw prompts, protected configurations or detailed internal records in the initial message.
What could make S2 unsuitable?+
A no-go result may follow when demand is almost entirely unique and necessary, acceptable output cannot be defined, pressure cannot be measured, boundaries are unclear or a controlled native comparison cannot be established.