From an F to a B: a Shopify performance programme that stuck

How a storefront went from a GTmetrix F to a B - WebP, lazy loading, load-on-demand, and the part nobody talks about: getting the team to keep it there.

Most performance work I have seen fails the same way. Someone runs Lighthouse, files a ticket called “improve page speed”, a developer spends two weeks on it, the score goes up, and six months later the site is slower than when they started.

The score is not the problem. The problem is that nothing in the team’s workflow changed, so the same weight comes back.

This is what we did on a Shopify Plus storefront, and - more usefully - what we changed so it did not come back.

Measure first, and pick one page

We started with a single template: the best-selling collection page. Not the homepage, which everyone stares at, but the page traffic actually lands on and converts from.

The starting numbers were not good:

  • GTmetrix grade F, 46%
  • Largest Contentful Paint 1.7s
  • Total Blocking Time 1.5s
  • Total page weight, one page load

Total Blocking Time was the number that told the real story. LCP at 1.7s is unremarkable. TBT at 1.5 seconds means the main thread was busy for a second and a half while the customer was tapping a filter and nothing happened.

One image, three jobsTHE ORIGINALPhone card300pxTablet700pxDesktop hero1400pxthe card was downloading the 1400 — sizes said 100vw
Getting `sizes` right cut more weight than any single conversion did.

Images: the cheap 60%

Images were the majority of the page weight, and almost none of it was necessary.

WebP for everything. Shopify will serve WebP through the image CDN if you ask it to. Most themes do not ask.

Lazy loading below the fold, eager above it. The rule sounds obvious and is constantly got wrong in both directions. Lazy-loading the hero image delays LCP. Eager-loading twelve product cards below the fold blocks everything.

Correct sizes. A product card 300px wide on a phone was downloading the 1200px variant, because the theme declared sizes="100vw" and moved on. Getting sizes right cut more weight than any single image conversion did.

Across the page this took 1.6MB out of the total.

BEFOREeverything, at oncefirst paint waitsAFTERthe pageCart drawerON TAPReviewsON SCROLLImage zoomON TAPnothing is removed — it just stops competing with first paint
The heavier win was ordering, not deleting.

JavaScript: load-on-demand

The heavier win was ordering. A storefront loads a lot of things it does not need yet - the cart drawer, the size guide modal, the review widget, the image zoom.

We moved those behind interaction. The cart drawer script loads when someone opens the cart, not when the page loads. The review widget loads when it scrolls into view. Nothing is removed, so nothing breaks; it just stops competing with first paint.

Third-party scripts got the same treatment where the vendor allowed it, and an audit where they did not. Some tags were still on the site for a campaign that ended a year earlier - the fastest optimisation available is deleting something.

Resource hints came last: preload the font and the LCP image, preconnect to the domains you genuinely hit early, and nothing else. A preload list of fifteen things is just a slower page with extra steps.

ONE PAGE LOADTotal Blocking TimeBEFORE1.5sAFTER63msPage weightBEFOREAFTER−1.6MB
The two numbers that changed how the page felt, not just how it scored.

The result

  • GTmetrix F (46%) → B (86%)
  • LCP 1.7s → 1.0s
  • TBT 1.5s → 63ms
  • −1.6MB total page weight

TBT is the number I would point at. 1.5 seconds to 63 milliseconds is the difference between a page that feels broken when you tap it and one that does not.

Presenting it to people who do not read waterfalls

Stakeholders do not care about TBT. What worked was showing the same page, side by side, on a throttled connection, recorded. Two videos. One is unusable for three seconds and the other is not. Nobody needed the numbers explained after that.

The same work on a store selling in thirty countries

The programme above is one storefront. The version of this problem I found harder was a brand running one Shopify theme across more than thirty markets.

Nothing about the techniques changes. What changes is that you cannot verify by looking. A script that a local team enabled for one market shows up in that market’s page weight and nowhere else, so a single audit tells you almost nothing. Regressions arrive market by market, and the person who caused one had no way of knowing they had.

Two things helped. Test on the market with the worst network and the busiest tag stack rather than the one head office reads. And make the expensive options harder to reach - if a section can only load a heavy widget through a setting that says what it costs, most people will pick the cheap path without being asked to care about performance.

Keeping it

This is the part that matters. We added a check to the review process: if a pull request adds a script or an image above a size threshold, that gets discussed before it merges. Not blocked - discussed. Most of the time there is a lighter way to do the same thing, and the person writing the code just did not know the budget existed.

A performance programme that is one developer’s project regresses. One that is part of how the team ships does not.