BenchmarkList's occupation map at benchmarklist.com sizes 66 job categories by economic output and colours each by how much of the role's work has direct AI benchmark evidence behind it. Software engineering: 528 benchmarks, 100% coverage, $288 billion in economic weight. Retail and sales professionals: 11 benchmarks, 60% coverage, $566 billion. That $566 billion figure is the largest single sector by economic weight on the entire map.

Those two rows tell most of the story.

What Benchmark Coverage Actually Measures

Benchmark coverage here is not the same as "AI can do this job." It measures how much of a role's work has formal AI evaluation evidence behind it: test sets, standardised tasks, scored outputs that let you compare one model against another on that specific kind of work. A role at 100% coverage has the full range of its task types represented across the benchmark corpus. A role at 17% coverage means most of what that job involves has never been formally evaluated at all.

Finance professionals: 87 benchmarks, 100% coverage, $424 billion. Customer service representatives: 71 benchmarks, 100% coverage, $124 billion. Corporate and legal professionals: 73 benchmarks, 100% coverage, $292 billion. These are the roles where a vendor's benchmark claim rests on substantial published test data you can actually look up.

The Roles That Are Almost Completely Untested

Event and food service management: 4 benchmarks, 21% coverage, $478 billion in economic weight. Production and manufacturing trades: 4 benchmarks, 34%, $398 billion. Construction and building trades: 3 benchmarks, 21%, $318 billion. Cashiers: 3 benchmarks, 17%, $119 billion. Postal workers: 0 benchmarks. Food production workers: 0 benchmarks.

These occupations represent hundreds of billions of dollars of economic activity. When a vendor approaches a business in any of these sectors and says their product is "benchmarked" and "proven," the published test base behind that claim ranges from thin to nonexistent.

What This Means for South African Retail, Hospitality, and Trades Businesses

South Africa's private-sector economy runs heavily on retail, food service, construction, and manufacturing. These are exactly the sectors with the lowest benchmark coverage on this map. A Gauteng retailer evaluating an AI tool for inventory management or customer communications is operating in territory where the research community has produced 11 benchmarks or fewer. The same vendor claiming "proven, benchmarked AI performance" to a software company is resting that claim on 528 evaluations. To your business, it rests on 11.

This does not mean AI tools for retail or hospitality cannot work. It means the vendor's benchmark citation is not the right evidence to ask for. The right question is: do you have performance data from businesses in my specific sector? Not from coding tasks or financial analysis evals that test nothing close to managing a shop floor or a restaurant kitchen.

How to Check Your Own Industry

The occupation map is at benchmarklist.com/research/occupation-map/. Find your SOC occupational category, check the benchmark count and coverage percentage. A vendor citing benchmark results for a sector with 3 or 4 evaluations behind it is making a fundamentally weaker claim than the same pitch sounds for a sector with 70 or 80. The count does not tell you whether an AI tool will work for your business. It tells you how much independent, standardised testing has been done on your kind of work. That is the honest starting point for evaluating what a vendor is actually selling you.