What is Throughput in Performance Testing?
Throughput is the amount of work an application completes over a given period. In performance testing, it is usually measured as requests per second (RPS ) or business transactions per second (TPS). Throughput tells you how much demand the system is processing; it does not, by itself, tell you whether the experience is fast, successful, or sustainable.
Throughput is one of the most useful and most misunderstood performance testing metrics.
A team may say that an application “handles 500 requests per second,” but that number only matters when you also know the response times, error rate, user journeys, and resource behavior behind it.
A healthy system can sustain the throughput that the business expects while meeting agreed response-time and reliability targets. When a bottleneck appears, throughput may stop rising, level off, or decline as errors, retries, and queues increase.
This guide explains throughput in plain language, shows how to calculate RPS and TPS, and explains what a throughput chart can reveal during a load test.
Throughput is the rate of completed work
In an application-performance test, throughput answers a simple question:
How much work did the system process per unit of time?
The unit depends on the level of the system and the question being asked.
Requests per second (RPS)
What it counts:
Technical requests received or completed by an application or API
Example: The product API processes 150 requests per second.
Transactions per second (TPS)
What it counts:
Completed business actions, which may include several requests
Example: The checkout service completes 12 paid orders per second.
Throughput over a period
What it counts:
Total work divided by elapsed time
Example: 9,000 requests in 60 seconds equals 150 RPS.
The formulas are straightforward:
RPS = total requests ÷ elapsed seconds
TPS = completed business transactions ÷ elapsed seconds
Before comparing results, define the counting rule. A load-testing tool may count all generated requests, while an application dashboard may count successful completed transactions. Both views can be useful, but they are not interchangeable. Report the throughput figure alongside the success or error rate so that nobody mistakes a high volume of failures for healthy capacity.
Microsoft defines RPS, or throughput, as the total requests that a load test generates per second and calculates it as the number of requests divided by total elapsed seconds. In k6, http_reqs counts the total HTTP requests generated by the test, while http_req_failed reports the failure rate. Together, these measures show volume and whether the application is actually handling it.
Let’s Imagine Throughput in the Real World

Let’s call our gas station Joe’s Gas. It has three pumps, and each pump takes one minute to fill a car. If three cars are at the pumps, Joe’s Gas can fill three cars per minute. That is the station’s throughput at that operating point.

This is Joe’s dilemma: no matter how many cars need gas, the maximum number that can be handled during a specific time frame will still be three cars per minute—until Joe’s Gas adds capacity or reduces the time needed to serve each car.

Now picture more cars entering the line. They have to wait, which creates a queue. The same thing happens in a web application. If an application receives 50 requests per second but can only complete 30 requests per second while meeting acceptable response-time targets, the other requests may wait, time out, be rejected, or trigger retries.
That is the important distinction: three cars per minute is the observed throughput at this operating point. Capacity is the highest throughput Joe’s Gas can sustain while still serving customers within an acceptable wait time. If only one car arrives in a minute, the actual throughput is one car per minute. If five cars arrive and only three can be served, the queue grows.
Throughput measures the flow of work. Capacity is the highest throughput a system can sustain while still meeting its defined service levels.
Automation Testing Free Courses
Throughput vs. concurrent users, response time, and latency
These metrics are related, but they are not synonyms.
| Metric |
The question it answers |
Why it matters |
| Concurrent users / virtual users |
How many users are active during the same test window? |
Defines the test population and workload driver. |
| Throughput (RPS/TPS) |
How much work reaches or is completed by the system per unit of time? |
Shows the demand and output created by that population. |
| Response time |
How long does a request take from sending it to receiving the complete response? |
Describes the experience and reveals degradation under load. |
| Latency |
How long until the first part of a response arrives? |
Helps isolate time spent before the response begins. |
| Error rate |
How often do requests or transactions fail? |
Prevents a high request volume from being mistaken for healthy behavior. |
| Capacity |
What throughput can the system sustain within its service objectives? |
Connects a technical test result to a release or business decision. |
Real Performance Testing Throughput Results
I use HP’s LoadRunner (which comes with a throughput monitor) for performance testing. But other tools like jMeter have similar meters.
In a typical test scenario, these would definitely happen:
- As users begin ramping up and making requests, the throughput created will increase as well.
- Once all users are logged in and processing in a steady state, the throughput will even out since the user load test stays relatively constant.
- If we wanted to find an environment’s throughput upper bound, we would continue increasing the number of users.
- Eventually, after a certain amount of users are added, the throughput will start to even out and may even drop.
- However, if the throughput enters this state, it is usually due to some kind of bottleneck in the application.
A fixed user count does not guarantee a fixed throughput. Ten virtual users with long think time may create less load than two users looping through an endpoint with no pause. Likewise, if the application slows down in a closed-loop test, each virtual user may complete fewer iterations, so throughput can fall even though the configured user count is unchanged.
Get TestGuild Partnership Plans
How to calculate throughput in a performance test
Start with a clear time window and total count.
If a test sends 9,000 API requests over 60 seconds:
9,000 requests ÷ 60 seconds = 150 RPS
If those API requests produce 720 successful checkout transactions during the same minute:
720 completed checkouts ÷ 60 seconds = 12 TPS
The distinction matters. One checkout transaction might require several API calls, including cart, inventory, payment, tax, order, and confirmation services. A request-level RPS measure is useful for infrastructure and service capacity. A transaction-level TPS measure is useful when the business asks whether customers can complete an outcome.
When you report the result, include the workload and quality signals in the same sentence. For example:
At the planned campaign mix of 500 concurrent virtual users, the checkout flow sustained 12 successful transactions per second. p95 checkout response time remained below two seconds, errors stayed below 0.5%, and database CPU remained below the agreed alert threshold.
That statement is much more useful than “the application handled 150 RPS.”
What a throughput chart should look like during a load test
A typical load test has three phases: ramp-up, steady state, and ramp-down.
During ramp-up, more virtual users begin their journeys, so throughput normally increases. During steady state, a well-sized system handling a stable workload should usually settle into a repeatable range. It does not need to form a perfectly flat line; small variation is normal because user journeys, response times, cache behavior, and data vary.
If you continue increasing load beyond the system’s sustainable capacity, throughput may level off or decline. This is the point where the system cannot turn additional incoming demand into additional successful work. Look for the symptoms that accompany the pattern: rising p95/p99 response times, growing queues, higher errors, exhausted connection pools, database waits, CPU saturation, or a slow downstream dependency.
Below are the LoadRunner throughput chart results for a 25-user test that I recently ran.
(Yes these screenshot are old (but still valid) I originally wrote this post in early 2012.)
Test #1
Notice that once all 25 concurrent users are logged in and doing work, the throughput stays fairly consistent. This is expected.

Test #2
Now notice what throughput looks like on a test that did not perform as well as in the last example.
All users login and start working; once all users are logged in and making requests, you would expect the throughput to flatline.
But in fact, we see it plummet. This is not good.

As I mentioned earlier, throughput behavior like the above example usually involves a bottleneck.
Work With Me
Test #3
By overlaying the throughput chart with an HP Diagnostics ‘J2EE – Transaction Time Spent in Element’ chart, we can see that bottleneck appears to be in the database layer:

In this particular test, requests were being processed by the web server. But in the back end, work was being queued up due to a database issue.
As additional requests were being sent, the back-end queue kept growing, and users’ response times increased.
To learn more about HP Diagnostics, check out how I configured LoadRunner to be able to get these metrics in my video: HP Diagnostics – How to Install and Configure a Java Probe with LoadRunner.
How to interpret a throughput drop
A throughput drop is a clue, not a diagnosis. It means the system is completing less work per unit of time than it was earlier in the test. The cause might be in the application, database, cache, queue, load generator, network path, test data, or a third-party dependency.
| Throughput pattern |
Common companion signals |
Questions to investigate |
| Throughput rises with the ramp and stays stable |
Response times, errors, and resources are stable. |
Is the observed rate sufficient for the planned demand and service objective? |
| Throughput plateaus while users increase |
p95/p99 latency rises; CPU, database, or queue utilization approaches saturation. |
Which resource is now the limiting constraint? Are requests waiting, timing out, or being rate-limited? |
| Throughput falls as load increases |
Error rate, retries, timeouts, queue depth, or garbage collection rises. |
Is completed work falling because a dependency, database, pool, or test generator is overloaded? |
| High throughput but high failure rate |
http_req_failed or transaction failures increase. |
Is the system processing useful successful work, or merely generating responses and errors quickly? |
The right answer comes from correlation. Do not stare at one throughput line and guess. Compare client-side throughput with request duration, error rates, logs, traces, infrastructure metrics, database waits, and queue depth over the same time interval.
TestGuild’s real example remains instructive: the web server was receiving requests, but back-end work queued at the database layer. As the queue grew, response times rose and throughput deteriorated. This is exactly why a performance test should observe the whole request path, not only one chart.
“It has detailed performance metrics including latency, requests per second, concurrency, and throughput.”
— Anvesh Malhotra, Performance Testing for Massive Scale
The practical lesson is that throughput belongs in a metric set. No single number can explain performance under load.
Free TestGuild Courses
What throughput is not
Throughput is not the same as network bandwidth. Network throughput describes data transferred across a network, often in bits per second. Application throughput describes requests or business transactions processed over time. Network conditions can affect application performance, but an article about application load testing should not treat SNMP, tcpdump, Wireshark, or packet size as the primary answer to an RPS/TPS question.
Throughput is also not the same as response time. A system may process a high volume of simple requests quickly while a small number of expensive journeys perform poorly. Conversely, increasing concurrency can raise RPS for a while even as response time starts to degrade. The test must specify the target for both flow and experience.
Throughput testing checklist
Before calling a result “good throughput,” confirm that the test defines the following:
- The work: Which endpoints and user journeys are in scope, and how is traffic distributed across them?
- The unit: Are you reporting generated requests, completed requests, or successful business transactions?
- The period: Which ramp-up and ramp-down periods are excluded from the steady-state analysis?
- The target: What RPS/TPS must be sustained, for how long, and at what expected growth margin?
- The quality bar: What are the acceptable response-time percentiles, error rate, and service-level objectives?
- The evidence: Which application, infrastructure, database, cache, and dependency metrics will explain a failure?
‘This is how throughput moves from a dashboard statistic to a trustworthy capacity result.
Frequently asked questions
Is throughput the same as RPS?
RPS is one common way to express throughput at the request level. Throughput can also be expressed as TPS when the unit is a business transaction, such as completed checkout, search, payment, or account-creation flow.
How do I calculate throughput?
Divide the count of requests or transactions by the elapsed time in seconds. For example, 9,000 requests completed in 60 seconds equals 150 RPS. Always state whether the figure includes failed requests and pair it with the relevant error rate.
Is high throughput always good?
No. High throughput is useful only if the application still meets its response-time, reliability, and resource objectives. A system can return large volumes of error responses or delay a critical customer journey while a dashboard still shows impressive request volume.
Why does throughput drop during a load test?
A drop often means the system has reached a constraint. Investigate response-time percentiles, error rates, queues, database waits, CPU or memory saturation, connection pools, third-party dependencies, and the health of the load generators before reaching a conclusion.
What is the difference between throughput and capacity?
Throughput is the work the system processes at a given time. Capacity is the maximum throughput it can sustain under a defined workload while meeting its agreed service levels.
Should I measure requests per second or transactions per second?
Measure both when they answer different questions. RPS helps describe technical load on services and infrastructure. TPS shows whether users complete meaningful business outcomes. Be explicit about the unit and avoid comparing them as though they were the same number.
The key takeaway
Throughput is the rate at which an application processes work. It is most useful when you connect it to a realistic user workload, RPS or TPS unit, response-time percentiles, success rate, and server-side evidence.
Do not ask only, “How much throughput did we get?” Ask: “What work did the system complete, at what rate, for how long, and did it remain fast and reliable while doing it?”