To load test a Magento 2, Adobe Commerce or Mage-OS store in 2026, pick a scripting tool for the traffic - k6, Gatling, Apache JMeter or Locust - and point it at a staging copy that holds an anonymised copy of your production data. For most stores I default to k6: the scripts are plain JavaScript a Magento developer can review, they live in git, and thresholds turn a run into a pass or fail in CI. The Magento Performance Toolkit still generates test data, but its JMeter scenario was removed from the codebase in 2022 and Adobe has published no replacement. Below is how the options compare, what is actually maintained on GitHub, and the Magento-specific traps that make load test numbers lie.
What does a Magento load test actually need?
Three parts, and most disappointing load tests skip one of them:
- Data - a catalogue, customers and orders shaped like yours. An anonymised copy of production is best; the Performance Toolkit's synthetic fixtures are the fallback.
- Traffic - a load generator that replays real journeys (browsing, search, layered navigation, add to cart, checkout) in roughly the proportions your analytics show.
- Observability - server metrics plus an APM or profiler during the test, so you see why response times climbed.
Size the load from your own numbers rather than round virtual-user counts: take peak-hour sessions and orders from analytics, add the headroom you want for a campaign, and express it as arrivals per second. k6 calls this the open model: new visitors keep arriving when the store slows down, which is exactly what happens on Black Friday. A fixed pool of virtual users (the closed model) quietly reduces the load as latency rises and can make a struggling store look stable.
Magento load testing tools compared
| Tool | Script language | Best for | Magento-specific notes | Licence / hosting |
|---|---|---|---|---|
| k6 | JavaScript (TypeScript is transpiled) | Journeys as code, pass or fail thresholds in CI | Cookie jar per virtual user handles sessions and carts; GraphQL is plain HTTP; a browser module adds front-end checks | AGPL-3.0; self-run, k6 Operator on Kubernetes, Grafana Cloud k6 |
| Gatling | Java, Kotlin or Scala; JavaScript and TypeScript SDK | Teams on Maven or Gradle, code-reviewed simulations | Recorder converts a HAR file of a real checkout into a simulation | Apache-2.0; self-run or Gatling Enterprise |
| Apache JMeter | Test plans built in a GUI, saved as XML; Groovy for logic | QA teams with existing JMX plans | The old 2.4.5 benchmark.jmx still runs; add an HTTP Cookie Manager for sessions | Apache-2.0; distributed remote mode, Azure Load Testing |
| Locust | Python | Python teams, complex user behaviour | HttpUser keeps cookies between requests; master and worker mode built in | MIT; self-run, Azure Load Testing |
| Artillery | YAML with JavaScript hooks | Quick scenarios, Playwright-based browser load | No maintained Magento scenarios that I could find | MPL-2.0; runs on AWS Lambda, AWS Fargate or Azure ACI |
| Siege | Command-line flags and URL lists | Hammering a list of URLs as a smoke test | Listed by Adobe for Cloud projects; not built for multi-step journeys | GPL-3.0; self-run |
All four main tools can drive a complete Magento journey, guest checkout included, because each keeps cookies per virtual user (k6 by default with a cookie jar per VU, JMeter once you add a Cookie Manager). None needs special GraphQL support - it is HTTP with a JSON body or a query string.
The real difference is who will maintain the scripts. JavaScript is the lowest bar for a Magento team that already writes front-end code, which is why I lean on k6. Gatling's Recorder and k6 Studio both turn a recorded browser session into a script. JMeter's latest release is still 5.6.3 from January 2024, though the project remains active. For load beyond one machine there are Grafana Cloud k6, Gatling Enterprise, JMeter's remote mode, Locust's workers and Azure Load Testing, which runs JMeter and Locust scripts only.
What happened to the Magento Performance Toolkit?
The toolkit is now half of what it was. The data side still ships with Magento 2.4.9 and Mage-OS: XML profiles in setup/performance-toolkit/profiles/ce (small, medium, medium_msite, large and extra_large) and bin/magento setup:performance:generate-fixtures, which writes a synthetic catalogue, customers and orders straight into the database. Adobe documents both in Generate data for performance testing.
The traffic side is gone. benchmark.jmx and benchmark_2015.jmx were deleted in August 2022 in commit ACPT-668, titled "Remove benchmark.jmx and benchmark_2015.jmx from magento-commerce/magento2* repos". They exist up to the 2.4.5 tags and its patch releases and are absent from 2.4.6 onwards. I looked for a successor - a separate Adobe repository, a Gatling or k6 port, a Composer package - and found none. Adobe's Application Testing Guide says its release load tests are "executed internally", so the scenarios Adobe uses are no longer public.
What the toolkit is still good for:
- A neutral baseline. Two environments loaded with the same profile hold the same data, so when you compare hosting, PHP versions or releases, data is not the variable.
- A reference scenario. You can still download
benchmark.jmxfrom the 2.4.5 tag and run it in JMeter's command-line mode against a fixture-loaded test store. It replays the Luma storefront of that era, so treat it as a reference for request sequences, not a model of your store.
Does Mage-OS have its own load testing tools?
Not as of October 2026. Mage-OS ships the same fixture profiles and generate-fixtures command as upstream, and none of the public repositories in the mage-os GitHub organisation is a load test suite or benchmark scenario. The Mage-OS developer docs have a general performance testing page that covers fixtures and names JMeter, Gatling and Locust.
The most useful Mage-OS material is a June 2026 benchmark post on mage-os.org comparing Mage-OS with Hyvä against Magento with Luma. Its method is worth copying: fixtures for identical data on both stacks, k6 for the traffic, and full-page cache switched off so PHP rendered every request. I could not find its scripts published.
Ready-made Magento load test scripts on GitHub
GitHub has dozens of Magento k6, Gatling, JMeter, Locust and Artillery repositories, almost all abandoned. As of October 2026:
- Genaker/magento-k6-performance - the most-starred k6 project for Magento, at about 20 stars. One k6 script that load-tests a single URL, with an option to append a unique query string so requests miss the full-page cache. No licence file; last commit May 2025.
- akadzik/gatling-performance-tests - parametrised Gatling (Scala) scenarios for Magento 2 browsing and cart. MIT; last commit August 2020, and its own to-do list notes that shipping and payment steps are missing.
- Older Gatling and k6 suites from 2013-2020, from Nexcess, Creatuity, EcomDev and others, target Magento 1 or 2.x releases long out of support, and are unmaintained or archived.
I found no maintained Locust, Artillery, Hyvä-specific or GraphQL load test suite for Magento. The practical conclusion: read these repositories for ideas, then write your own scripts. A store's URLs, extensions and checkout are specific enough that a generic script would mislead you anyway.
Which option fits which situation?
| Situation | What I would use | Why |
|---|---|---|
| Capacity test before launch or peak season | k6 scripts of your own journeys, anonymised production data | Realistic data and traffic; scripts are reusable in CI |
| Comparing hosting, PHP versions or releases | Toolkit fixtures plus the same k6 script on both environments | Keeps data out of the comparison |
| QA team already maintains JMX plans | JMeter in command-line mode | Reuse what the team knows; avoid the GUI for the load itself |
| JVM-based team, or you want to record a checkout | Gatling | Recorder plus simulations as code |
| Python-based team | Locust | Behaviour as regular Python |
| Headless or PWA storefront | k6 or Gatling with the storefront's real GraphQL calls | Same queries, same HTTP method, same cache behaviour |
| Adobe Commerce on cloud | Any of the above, plus a support ticket and New Relic | Adobe asks to be told; New Relic is included |
| Is the store fast for real visitors? | CrUX field data, not a load test | Measures what real Chrome users experience |
| Why does it slow down under load? | Blackfire, Tideways, New Relic or OpenTelemetry during the test | Load tools show the symptom, profilers the cause |
What I use in Magento performance audits
In a Magento performance audit I work on a copy of your code and an anonymised copy of the production database, and nothing runs against your live store. My default setup:
- Data: the anonymised production copy on an environment sized like production. Fixtures only when I need a neutral baseline, for example to separate a hosting problem from a code problem.
- Traffic: k6, with three scenario types run separately and then together - cacheable browsing (home, category, product), low-cache work (search and filtered category pages, which create many unique URLs) and cart plus checkout.
- Observability: CPU, PHP-FPM busy workers, MySQL, Redis and the Varnish hit rate, plus a profiler on sampled requests.
A minimal browsing scenario: an open-model ramp, think time between pages and a pass or fail budget.
import http from 'k6/http';
import { check, sleep } from 'k6';
const BASE = __ENV.BASE_URL; // a staging copy, never production
const PAGES = ['/', '/your-category.html', '/your-product.html'];
export const options = {
scenarios: {
browse: {
executor: 'ramping-arrival-rate', // open model
startRate: 1,
timeUnit: '1s',
preAllocatedVUs: 50,
maxVUs: 300,
stages: [
{ target: 5, duration: '5m' },
{ target: 5, duration: '10m' },
],
},
},
thresholds: {
http_req_failed: ['rate<0.01'],
http_req_duration: ['p(95)<1500'], // example budget - set your own
},
};
export default function () {
PAGES.forEach((path, i) => {
if (i > 0) sleep(2 + Math.random() * 3); // think time between pages
const res = http.get(BASE + path, { tags: { name: path } });
check(res, { 'status is 200': (r) => r.status === 200 });
});
}
Run it with k6 run -e BASE_URL=https://staging.example.com browse.js, raise the arrival rate step by step between runs, and look for the knee where throughput stops rising and p95 starts climbing. The cart and checkout scenarios are where the real work is, for the reasons in the next section. For the release-level checks that sit around a load test, see the Magento testing checklist.
Magento-specific pitfalls that skew load test results
- Fixtures are not your store. They have uniform attributes and none of your extensions or custom code, which is where most of the slowness I find comes from.
- Cache hit ratio decides the result. Cacheable guest pages served by Varnish or Fastly mostly measure the cache. Report warm-cache and cache-miss runs separately and watch the hit rate during the test; cart, checkout and customer account pages always reach PHP.
- Form keys fail quietly. Every storefront POST needs a
form_keythat matches the session. On cached pages, Magento's form-key-provider.js keeps it in aform_keycookie (generating one if missing) and copies it into the forms, and a server-side plugin registers the cookie value with the session, so your script must set that cookie and post the same value. With a bad key, add to cart returns a normal-looking redirect and a "Your session has expired" message, so a status-code check reports success. - Checkout differs per frontend. Luma's checkout calls REST endpoints. Hyvä Checkout is built on Magewire, a PHP-driven component framework, so checkout interactions are server round trips your script has to replay. Headless storefronts use GraphQL mutations such as
addProductsToCartandplaceOrder. Record your real checkout and script that, with an offline payment method such as Check / Money order, never a live gateway. - GraphQL caching depends on the HTTP method. Adobe's GraphQL caching docs state that queries sent in a POST body are not cached; only GET requests with the query in the URL are. Match what your storefront sends.
- Admin 2FA blocks scripted admin logins. Two-factor authentication is on by default in 2.4.x. On a disposable environment only, disable it with
bin/magento module:disable Magento_AdminAdobeImsTwoFactorAuth Magento_TwoFactorAuth(2.4.6 and later). - Same-host and parity bias. Generate load from a separate machine, ideally in your customers' region. Match production on PHP version and OPcache, PHP-FPM pool sizes, Redis, Varnish, search engine and database size, or the result describes the staging box.
- Adobe Commerce on cloud has a process. Adobe's staging and production testing guide asks you to open a support ticket before testing, stating the environments, tools and time frame, and to update it with results.
Profiling and APM: finding out why it slowed down
A load generator tells you that p95 climbed at a certain arrival rate; it does not tell you which observer, query or extension caused it. Run one of these alongside the test:
- Blackfire - a PHP profiler with documented Magento and Adobe Commerce Cloud integrations. Ideal for comparing one request before and after a change.
- New Relic APM - included with every Adobe Commerce on cloud project, Pro and Starter, with infrastructure and log monitoring on Pro.
- Tideways - PHP monitoring, profiling and tracing with a dedicated Magento 2 offering.
- OpenTelemetry PHP - the vendor-neutral route. Traces, metrics and logs are all marked stable, with zero-code instrumentation through a PECL extension or Linux packages.
My order of work: profile single requests on a quiet environment and fix the obvious, then load test with APM on, then profile again under load. The causes I find most often are in Why is my Magento store slow?.
Load tests vs real-user data: what CrUX tells you
Load tests answer one question: how much traffic the back end takes before it degrades. They say little about how fast the store feels. For that you need field data: the Chrome UX Report (CrUX), shown in PageSpeed Insights and Search Console, aggregates Core Web Vitals from real Chrome users over a 28-day collection period. Google's thresholds for a good experience are LCP within 2.5 seconds, INP of 200 milliseconds or less and CLS of 0.1 or less, assessed at the 75th percentile.
A store can pass a load test and fail Core Web Vitals because of front-end weight, and the reverse can happen too. Track both, and use a synthetic browser run - Lighthouse, or the k6 browser module - as a lab check between CrUX updates.
Load test results you do not trust, or a store that slows down under traffic for no clear reason, are exactly what my performance audit covers.
