Skip to main content
Zimei TechnologyEnterprise AI · Development and delivery
EnglishEN
Discuss a project

Enterprise AI Project Wiki · Requirements and business scenarios

Performance requirement

Also calledPerformance requirements · Performance and timing requirement · System performance requirement · Performance specification

Definition

A performance requirement is a normative quantitative statement of the speed, latency, deadline, throughput, concurrency, processing scale, capacity, or resource-use level with which a named system object must perform a named function. It also defines workload, data, environment, measurement boundary, statistic, threshold, and allowed failure so that the result is reproducible, verifiable, allocatable, traceable, and change-controlled.

A performance requirement says how fast, how much, and how efficiently—not that the system should be fast

Performance attaches to a function and boundary. A page shall open in two seconds omits which page, start and end points, backend and network scope, data and load, and whether an average or percentile decides compliance. Development, test, customer, and monitoring can otherwise each prove a different result.

Performance may govern end-to-end interaction response, event latency between points, a batch deadline, successful work per unit time, simultaneously active actors, maximum data or queue scale, or resources per unit of work. Metrics do not substitute for one another: more concurrency can slow each request, and a low mean can hide damaging tail waits.

NASA distinguishes what function is performed from quantitatively how well it is performed and recommends a minimum acceptable threshold alongside the desired baseline. Too loose fails the mission; unnecessarily tight tolerances exclude feasible alternatives and add cost without value.

What common performance metrics measure

Response time, throughput, concurrency, and resource utilisation often share one table but cannot replace one another. A fast page does not prove peak capacity, and spare server capacity does not prove that a user received a usable result in time. Knowing the question behind each measure prevents an easy metric from standing in for the real task.

Response time

Duration from a defined request or user action to an agreed usable result, stating whether client, network, queueing, retries, rendering, and asynchronous completion are included.

Latency

Time for an event between two observation points in an interface, message stream, freshness path, or pipeline, including clock synchronisation and buffering assumptions.

Throughput

Successful transactions, messages, documents, tokens, or data per unit time, reported with failure, refusal, and backlog so dropping work cannot create apparent throughput.

Concurrency

Users, sessions, tasks, or transactions connected, active, queued, or executing simultaneously. Select the definition that reflects business behaviour and resource contention.

Capacity

Maximum data, user, task, storage, queue, or call scale while response, error, and resource boundaries still hold—not an isolated maximum count.

Deadline and period

A task completes before a business or control deadline or starts at a specified frequency. Missing a hard real-time deadline differs from slower ordinary interaction.

Resource efficiency

CPU, memory, storage, network, accelerator, energy, or external-call consumption for a stated workload, classified as limit, mean, peak, or unit-of-work cost.

Accuracy and precision

Some systems-engineering taxonomies include measurement accuracy and precision as performance while software quality may classify correctness separately. Declare the scheme; speed cannot excuse wrong results.

Queue, backlog, and loss

Queue depth, wait, expiry, discard, backpressure, and drain time reveal deferred work and can expose capacity risk earlier than average response.

A threshold has no shared meaning without a workload model

“Complete within two seconds” needs a user population, request size, concurrency, and dependency state. Otherwise a short request can prove success in development while production receives long documents at month-end peak. The workload model connects the test condition to real business volume.

Business transaction mix

Define proportions and arrival patterns for reads, writes, uploads, generation, exports, and batches. Testing the lightest endpoint does not represent the service.

User and session behaviour

Separate registered, connected, active, thinking, bursting, and executing populations and model the function mix of each role.

Data scale and distribution

Record count, file size, field complexity, history, index state, hot and cold data, language, and exceptional but valid inputs rather than an average object alone.

Arrival rate, peak, and duration

Distinguish normal, seasonal peak, instantaneous burst, sustained peak, and growth assumption. A queue can absorb a burst but not substitute for sustained capacity.

System configuration and dependencies

Fix application, database, cache, network, region, instances, quotas, third-party and model versions, autoscaling, warm-up, and rate-limiting conditions.

Initial state and data preparation

State cache warmth, pools, queues, indexes, background jobs, accounts, and data. Reporting only a warmed best case misrepresents first-use and post-failure performance.

Success, refusal, and degradation

Success includes functional correctness. Count timeout, 5xx, business refusal, throttling, fallback, reduced-quality output, and asynchronous acceptance separately.

Representativeness and expiry

Record the business evidence, populations and periods covered, unknowns, and the demand or behaviour change that requires a new model.

What to record in a verifiable performance requirement

A performance requirement is a measurement contract. It identifies the start event, the result that ends timing, load and data conditions, and whether a percentile, maximum, or another statistic determines success. Omit one and the result becomes difficult to reproduce.

Stable ID, object, and function

Name the end-to-end service, interface, job, or component, its function and version. Component compliance does not prove the user path.

Load, data, and duration

Reference a governed workload with concurrency, arrival, transaction mix, data scale, peak duration, and growth boundary.

Environment and dependency state

Specify topology, resources, region, network, cache, supplier, background work, and normal or degraded state, disclosing differences from production.

Measurement start, end, and boundary

Define observable start and completion and whether queueing, retries, client, network, tools, and asynchronous follow-up are included.

Formula, unit, and statistic

Specify milliseconds, work per second, or resource per transaction; percentile, maximum, mean, distribution, or confidence; aggregation window and minimum sample.

Threshold, objective, and margin

Separate minimum threshold, desired baseline, engineering margin, and capacity warning, including one-event, ratio, or consecutive-window judgement.

Error, timeout, and exclusion

Include failure, throttle, refusal, degradation, cancellation, and timeout, with authorised exclusions so the slowest failed samples are not silently deleted.

Source, rationale, and consequence

Trace to user tolerance, business deadline, upstream rate, interface limit, risk, cost, or regulation and state the consequence of failure.

Verification, monitoring, and authority

Name tools, environment, data, repetitions, witnesses, evidence storage, production metric, owners, and review triggers.

Allocation, version, and state

Connect function, architecture, capacity model, supplier commitment, tests, and acceptance and record TBD, approval, waiver, alternative, and change.

How to derive evidence-based thresholds

A defensible threshold does not come from “as fast as possible.” Start with the consequence of user waiting or a missed batch deadline, then combine current measurements, growth, dependency capacity, and cost. If evidence is not ready, approve a measurement phase rather than inventing a permanent number.

  1. Identify the business outcome

    Find user actions, deadlines, upstream interfaces, and batches sensitive to wait, backlog, and scale, and the actual consequence of insufficient performance.

  2. Collect baseline and growth evidence

    Use production logs, forecasts, seasonal peaks, research, incidents, interface limits, and supplier data, recording uncertainty. Measure a baseline before inventing one.

  3. Build tiered workload models

    Define normal, peak, burst, batch, exceptional, and degraded conditions with mixes, distributions, duration, and future growth—not one concurrent-user count.

  4. Define end-to-end measurement

    Begin with the business-visible start and end, then decompose client, network, application, database, model, tool, and queue stages as needed.

  5. Set threshold and desired baseline

    Derive a minimum and objective from tolerance, dependency, risk, and current evidence, testing feasibility and retaining useful design trade space.

  6. Analyse performance, quality, and cost

    Compare latency, throughput, correctness, security checks, reliability, resources, cloud and model cost, complexity, and supplier lock-in with residual risk.

  7. Approve evidence and margin

    Business approves timing and scale, engineering measurement and feasibility, operations peaks and alerts, and procurement supplier quotas; jointly approve margin and failure handling.

  8. Baseline, monitor, and revisit

    Link scripts, results, monitoring, and capacity plan; remeasure after material demand, data, architecture, model, interface, or hardware change.

Why averages are insufficient

A one-second average may combine half a second for most people with twenty seconds for a few. If the slow group contains large documents, high-value customers, or month-end tasks, the attractive average is irrelevant. Examine distributions, tails, failures, and distinct business paths.

Means hide tail latency

Many fast requests dilute a few very slow ones that may affect critical users or large files. Report median, approved high percentiles, and the distribution.

Percentiles need an aggregation scope

A p95 or p99 names endpoint, role, region, version, and window. Aggregating samples then calculating differs from averaging per-instance percentiles.

Maximum needs severe-failure treatment

Maximum is noise- and sample-sensitive, but one missed business or control deadline may still be unacceptable and needs a hard-failure rule.

Success-only samples create false compliance

Calculating only returned successes excludes timeouts and errors. Response distribution and failure, cancel, throttle, and discard must describe one population.

Coordinated omission hides queueing

Waiting for each response before sending the next automatically reduces pressure when the system slows. The generator must represent the real arrival model.

Short tests do not prove endurance

Warm-up, cache, garbage collection, leaks, backlog, exhaustion, and supplier quotas may appear only over time. Duration must match peak and risk.

Repeat and describe uncertainty

Shared infrastructure, networks, and model services vary. Record repetitions, sample size, distribution, noise, and confidence rather than selecting the best run.

How to verify performance and carry it into production

A laboratory result is meaningful only when environment, data, dependencies, and load represent intended operation closely enough. Production monitoring then reuses the same timing boundary and segmentation; otherwise the two-second test and the two-second dashboard may measure different events.

Build a reproducible environment

Lock artefact, configuration, resources, network, dependencies, data, cache, and generator and record how production differs and in which direction that may bias results.

Calibrate load and instruments

Ensure the generator is not the bottleneck, clocks, sampling, traces, and resource monitors are accurate, and transactions are functionally successful.

Run risk-selected test types

Use baseline, load, peak, stress, endurance, capacity, and degraded tests as required. One peak run does not establish every performance property.

Observe the system and dependencies

Correlate end-to-end and stage time, throughput, queue, errors, CPU, memory, storage, network, database, supplier service, and cost to locate the bottleneck.

Retain versioned raw evidence

Store requirement, workload, environment, scripts, raw results, analysis, deviations, and reruns—not only a screenshot or passed statement.

Use approved criteria for acceptance

Name formal version, representative data, environment, window, and pass rule. Any environment difference and residual risk need authorised acceptance.

Map consistent metrics to production

Use matching boundaries and strata for trends, budgets, alerts, and authority. Production traffic does not replace controlled verification or end continuing observation.

Constructed performance requirements for a procurement-request service

This customer-independent procurement service turns interactive waiting, batch processing, and peak capacity into complete measurement conditions. Values remain variables because targets require actual load, business tolerance, architecture, and cost evidence rather than an industry marketing figure.

PERF-101: interactive check response

Under workload W1, data D1, and pre-production E1, from an authorised applicant's request until the page receives all missing items and rule IDs, successful end-to-end response shall meet p95≤T1 and p99≤T2, with timeout and system-error ratio≤F1.

PERF-102: sustained submission throughput

During peak mix W2 for H1, the service shall complete at least Q1 formal submissions per minute while preserving PERF-101, state integrity, and duplicate prevention; queued-after-test and dropped requests do not count.

PERF-103: concurrency and queue

With no more than C1 active submission transactions and dependencies at approved levels, p95 queue wait shall be≤T3; above C1 the system applies approved backpressure or refusal with an identifiable response.

PERF-104: audit export deadline

For authorised data set D2, export shall complete with integrity evidence before business deadline B1 without causing the interactive path to violate PERF-101.

PERF-105: resource efficiency

Under W1, CPU, memory, network, model tokens, and external-call cost for one successful check shall be recorded under formula M1 and meet limit U1 approved by capacity and cost review.

PERF-106: cold start and scaling

From approved zero or low capacity when burst W3 arrives, the system shall reach Q2 within T4 while requests follow defined queue, degrade, or refusal rules; cold samples are not excluded.

Enterprise AI performance covers the whole task chain and variable cost

Model generation is only one part of user waiting. Access checks, retrieval, reranking, tools, content validation, and human queues add time and may determine whether the result is usable. Measure the end-to-end task, quality, and cost together so speed does not quietly remove evidence or controls.

Decompose without losing end to end

Measure access, context assembly, retrieval, reranking, model queue and generation, tools, post-processing, and transfer while retaining request-to-usable-result time.

Text appearing quickly does not mean the work is complete

Streaming can reduce perceived waiting, so first usable feedback is worth measuring. The user often needs a complete, validated, and permitted result, whose time must be measured separately. Optimising the first visible word can otherwise hide later tool and validation delays.

Model input and output scale

Capture distributions of prompt, context, documents, images, output tokens, tool count, and history, including long-tail valid tasks rather than short-prompt means.

Judge quality, latency, and cost together

Less retrieval or reasoning can improve speed and cost while reducing correctness. Performance uses configurations meeting approved task quality and reports severe failure.

Cover supplier quotas and variation

Model regional network, provider queueing, rate limit, burst quota, version change, and degradation with timeout, retry, switch, queue, and refusal rules.

Tools and human queues are task time

For database, API, review, or approval tasks, state automatic-stage and end-to-end business deadlines. Generation time is not business completion.

Cache and batch within knowledge controls

Reuse and batching require permission, version, freshness, invalidation, and isolation so speed does not come from stale or cross-user content.

Monitor each composite and path

Stratify latency, throughput, errors, tokens, and cost by model, prompt, retrieval, tool, language, task, and fallback; a global mean hides a degraded risky path.

Supplier change triggers regression

Model version, context limit, price, quota, region, or tool changes require representative load, quality, and cost reevaluation and an updated capacity decision.

How performance requirements differ from neighbouring concepts

Performance, availability, reliability, capacity, and service levels interact but use different events and decisions. A performance requirement governs time, volume, and resource behaviour for a task under a stated load. Deployment size or a supplier quota influences it but does not prove the user's task meets the target.

ConceptBoundary from a performance requirement
Functional requirementStates what the system does and produces; performance quantifies how quickly, how much, or how efficiently under named conditions. Correctness is a precondition for valid samples.
Non-functional requirementThe wider quality or constraint set may include security, reliability, and maintainability; performance focuses on quantitative time, throughput, capacity, and resources.
Response time and latencyMetrics, not complete requirements. Function, boundary, load, environment, statistic, threshold, and error rules make them verifiable obligations.
Concurrency and capacityConcurrency describes simultaneous activity; capacity is supportable scale while other boundaries hold. An undefined supported-user count proves neither.
ScalabilityConcerns retaining behaviour as load changes by adjusting resources or architecture; performance gives concrete loads and outcomes, while scalability also covers scaling time, efficiency, and limits.
Reliability and availabilityReliability concerns correct service over time and availability the proportion of usable state. Fast successes cannot hide widespread failure or unavailability.
Recovery time objectiveTargets restoration after disruption in continuity governance. It is temporal but not ordinary request performance.
Performance testingProduces evidence under controlled load and environment; the requirement is the normative target. Tool defaults cannot create the need.
Monitoring metricObserves operations; the requirement provides the target and conditions. Measurable does not mean justified or already verified.
SLI, SLO, and SLASLI defines a service indicator, SLO its target, and SLA responsibilities and remedies. They may use performance measures but differ in lifecycle, exceptions, authority, and legal effect.
Performance budgetAllocates an end-to-end target across resources, components, or stages for design control. It derives from the parent requirement and still needs whole-path verification.
Cost requirementGoverns budget or unit economics. Resource efficiency influences cost, but prices and commercial terms change; CPU or token use alone is not the complete cost obligation.

Sources and scope

This entry explains a performance requirement in systems and software requirements engineering: a quantitative obligation for time behaviour, throughput, concurrency, capacity, processing scale, or resource efficiency while a system performs named functions under explicit load, data, environment, and time-window conditions. It answers how quickly, how much, with what resources, or by which deadline, and is commonly a subtype of quality or non-functional requirements. It does not fully define functional or non-functional requirements, reliability, availability, recovery-time objectives, scalability, capacity planning, SLI/SLO/SLA, performance testing, monitoring, cost optimisation, model accuracy, or user experience. Related concepts still need distinct objects, authorities, and decisions. No latency, concurrency, or cost value is universal; authorised parties approve it from business, risk, environment, and experimental evidence.