Original researchMeasured, dated, citable
Numbers we measured ourselves, free to cite with their dates.
Writing about Chinese suppliers, supply-chain risk or company registration usually means repeating someone else's assertion. This page collects the measurements we produced instead: each study states its frame, method, query date and what the numbers cannot show. If you are writing an article, an entry or a paper, this is the index.
Ten original studies
- Can you find a Chinese supplier from its English name? (measured 12 August 2026) — 106 Chinese manufacturers taken from the US NHTSA vPIC register, each searched by the registered English name that register holds: 57.5% returned a candidate, and for 77.0% of those the top candidate’s own registration record carried the same English name — an end-to-end confirmable rate of 44.3%. On the 22 companies carrying both, brand or short names returned far more often (90.9% against 59.1%) and were mostly the wrong company.
- 264 top Chinese manufacturers: registration study (queried 8 August 2026) — a census, not a sample: every company on three official provincial excerpts of MIIT’s “Little Giant” list, read for scope wording, former names, capital fields and status. Findings: 74.6% carry import-export wording (so 25.4% do not), 53.4% have changed legal name, 37% show paid-in below subscribed capital, all 264 resolved as live registrations.
- How often Chinese company records change (queried 8 August 2026) — the change logs of the same 264 companies: 54.9% amended their registration within 12 months, 78.4% within 24; median lifetime change count 39; median days since last change 320.
- China official verification sources: availability panel (measured 8 August 2026) — eight official hosts requested three rounds each from inside mainland China with same-session controls: five refused their front page; none of the three that loaded carried an English-version marker; one court host served a browser and refused a command-line client.
- How much does a check digit protect you? (measured 9 August 2026) — 2.9 million deliberate transcription errors run through the USCI (GB 32100-2015) and VIN (49 CFR 565) check algorithms with a fixed seed: USCI caught 100% of single errors and 100% of data-position swaps; VIN caught 92.5% and 88.1%, with every single-substitution miss traced to same-value transliteration.
- How many Chinese trailer makers are NHTSA-registered? (captured 9 August 2026) — an exhaustive count of NHTSA's public vPIC manufacturer database: 22,881 registrants across 92 countries, 2,589 from China, and 106 Chinese trailer manufacturers — the full 106-name list published as CSV.
- Certificate registries: who can actually open them? (measured 9 August 2026) — eight certificate-verification registries probed three rounds each over two network paths with same-session controls: 5 of 8 answered a mainland-China connection; UL's Product iQ refused a scripted client on every path; both US government hosts produced no response at all through the proxy egress — while on the direct path one served (NHTSA vPIC, 200) and the other refused (FCC, 403), which are different facts.
- China company registry availability (measured 5 and 8 August 2026) — the GSXT question specifically, with control hosts, a user-agent pair experiment, an independent same-method replication three days later — and one published conclusion we withdrew, documented on the page.
- What a China FDA registration tells you: 4,973 measured (source export dated 10 August 2026) — all 41,745 China listing records in openFDA deduplicated to 4,973 establishments, with declared roles, name continuity and agent concentration reported as aggregate data.
- One model translated, another checked: 11 of 98 wrong (reviewed 12 August 2026) — 98 interface strings across 15 languages on this site, back-translated by a second, independent model before shipping: 11.2% came back wrong, and the four automated string checks that ran first caught none of them. The most expensive one told Arabic-reading buyers to ask a Chinese factory for a work permit instead of a business licence.
Original measurement106 companies
Can you find a Chinese supplier from its English name?
Buyers outside China usually hold an English name from an email signature, catalogue or platform storefront, not the Chinese registered name. Whether that is enough to find the company is a rate, not a yes-or-no promise. We measured it on 12 August 2026.
English-name lookup: the numbers
The frame is a census, not a sample: every Chinese manufacturer we had already enumerated in the US NHTSA vPIC register, searched by the English name that register holds for it.
| Measure | Result | Of |
|---|---|---|
| Returned at least one candidate | 57.5% | 61 of 106 |
| — top candidate has an English name on its registration record | 83.6% | 51 of 61 |
| — that recorded English name equals the input | 77.0% | 47 of 61 |
| — top candidate carries an 18-character Unified Social Credit Code | 83.6% | 51 of 61 |
| End to end: input an English name, get a confirmable match | 44.3% | 47 of 106 |
A search returning something is not the same as a search returning the right thing. When the entity you land on independently records the English name you started from, you have a checkable link between the two. Otherwise you have a guess with a company name attached.
A candidate without an 18-character code cannot be taken to a dated registration check, so it is a dead end regardless of how right it looks.
The brand-name trap
Twenty-two companies in the frame carry both a full registered English name and a brand or short name. Searching the same 22 companies both ways:
| Input | Returned a candidate |
|---|---|
| Full registered English name | 59.1% (13 of 22) |
| Brand or short name | 90.9% (20 of 22) |
The higher number is the worse one. Inspecting the brand-name results one by one, most were a different company: a three-letter brand returned three Hong Kong shell companies sharing those letters while the actual manufacturer — a specialist vehicle maker in Hubei — was not among them. Another brand returned a machinery firm in Guangdong when the company sought was a machinery firm in Henan with a homophone name.
The cause is constraint count. A full registered English name usually encodes a place, a company style and an industry word. A brand name encodes only the style, and Chinese company styles repeat heavily across provinces and industries.
This is a general search-interface hazard: a query that returns more is not a query that works better. The measurement that matters is the share of returns you can independently check.
What this means at a desk
- An English name is a starting point for about two suppliers in five. That is worth trying, but a plan that assumes it will work fails three times in five.
- Never take the first result. Ranking is relevance, not identity. In our brand-name runs the top result was frequently unrelated and presented exactly as convincingly as a correct one.
- Ask for the business licence (营业执照). It carries the registered Chinese name and 18-character code, turning a search problem into a lookup.
- A refusal is evidence too. A supplier unwilling to provide a business-licence photo has not proved fraud, but it has blocked the shortest route to confirming the entity and should not be silently treated as a clean result.
- A name match is not a supplier check. It establishes which entity you are discussing. Active status and scope live in the registration record; whether the person emailing you is connected to it lives in no registry.
Method, and what this study cannot say
Each company's registered English name was submitted to a licensed Chinese business-information platform from a mainland network egress. The platform refused requests from outside China during this measurement, so a foreign desk cannot reproduce the run merely by copying the query. For every search that returned candidates, the top candidate's registration record was retrieved and its recorded English name compared with the input after normalisation: uppercase with all non-alphanumeric characters removed, so that Co., Ltd. and CO.,LTD compare equal.
Two independent runs were executed seven hours apart and returned identical counts on every metric. That is evidence of a stable index, not of correctness.
What it cannot say
- It is one sector. The frame is vehicle and trailer manufacturers registered with a US federal authority. Every company had a reason to record an English name somewhere, so this is plausibly an upper bound for Chinese exporters generally, not an average.
- The paired comparison is 22 companies. The direction — brand names return more and are mostly wrong — is clear and mechanically explicable. The specific percentages are not stable at that size.
- vPIC English names are not Alibaba storefront names. An English name recorded with a US federal regulator is more likely to be the company's registered English name; a storefront name is chosen for marketing and can change. Buyers holding only a storefront name should expect to do worse than 44.3%, not better.
- A matching English name confirms nothing else. Status, scope, export capability and counterparty identity are separate checks.
- One platform, one date. Coverage differs between providers and changes over time. The measurement date travels with the numbers.
English-name lookup questions
Can I find a Chinese company using only its English name?
Often, but not reliably. Across 106 Chinese manufacturers taken from the US NHTSA vPIC register, searching by the registered English name returned at least one candidate for 57.5%. For 77.0% of those returns the top candidate's own registration record carried an English name equal to the input, which is what makes a match checkable rather than plausible. End to end that is 44.3% — so roughly two in five English names lead to something you can confirm, and the rest need the Chinese registered name or the business licence.
Why do brand names return more results but help less?
On the 22 companies in our frame that carry both a full registered English name and a brand or short name, the brand name returned a candidate 90.9% of the time against 59.1% for the full name — and most of those extra returns were the wrong company. A full registered English name usually carries three constraints at once: a place, a company style and an industry word. A brand name carries only the style, and Chinese company styles repeat heavily across the country, so the search matches something with the same syllables in a different province and a different industry.
What does a matching English name actually prove?
That the entity you found records the same English name you were given. It does not prove the company is active, that its scope covers your product, that it can legally export, or that the party emailing you is that company. Those live in the dated registration record and, for the last one, nowhere in any registry. The English name match is a way to stop guessing which candidate to check, not a substitute for checking.
Is 44.3% good or bad?
It is a sector-specific reference point, plausibly an upper bound rather than an average. Our frame is vehicle and trailer manufacturers registered with a US federal authority, which means every company in it had a reason to record an English name somewhere. Sectors that export less, or sell only through platforms, would plausibly do worse. The number's practical use is comparative: it says an English name is a usable starting point for about two in five suppliers, and that the remaining three in five are not a search problem but a document problem.
Citing the English-name study
You may quote or reproduce these figures, including commercially, provided the measurement date, frame and limits travel with them. Machine-readable metrics with numerators and denominators are in the result file. The frame comes from our census of Chinese manufacturers in NHTSA vPIC, which readers can re-derive from the vPIC API.
Original measurement98 strings · 15 languages
One model translated. Another one checked.
We localised this supplier-verification tool into fourteen languages we do not read. Before shipping, a second, independent model back-translated every string so we could see what our own words actually said. This is what it found, including the one that would have sent buyers to ask for the wrong document.
Machine-translation review: the numbers
| Measure | Result | Of |
|---|---|---|
| Strings flagged as wrong | 11.2% | 11 of 98 |
| — wrong domain term | 45.5% | 5 of 11 |
| — grammar, voice or collocation | 27.3% | 3 of 11 |
| — meaning close but not exact | 27.3% | 3 of 11 |
| Would have sent the reader to the wrong document | 9.1% | 1 of 11 |
| Flags our automated checks caught first | 0% | 0 of 11 |
Flags fell in six of the fourteen non-English languages — Arabic, Vietnamese, Hungarian, Russian, Polish and Portuguese — with one to three each. The other eight came back clean. We are not drawing a conclusion from that split. Eleven flags spread across six languages is far too thin to say which languages a model handles worse; at this size any pattern could be noise, and saying otherwise would be the same overreach this study is about.
The worst one
One string tells a buyer what to do when the tool finds nothing: ask the supplier for a photo of the business licence (营业执照) — the Chinese registration document that carries the registered company name and the 18-character Unified Social Credit Code.
The Arabic version used a phrase meaning work permit.
It was grammatical. It used the right register. It passed the character-set check, the forbidden-word list and the length budget. A buyer reading it would have asked a Chinese factory for an employment document — and when that produced nothing useful, would have concluded the supplier was being evasive rather than that our sentence was wrong. The failure would have been silent and misattributed.
Two more of the same kind: a Hungarian gloss that meant an operating permit for premises, and a Polish phrase that was not the term for that document at all. Each one was a plausible-looking word in the right semantic neighbourhood. That is exactly what makes them expensive — a translation that is obviously broken gets fixed; a translation that is confidently wrong ships.
Why the automated gates missed all eleven
Before the review, every string had to pass four checks. Each corresponds to a real failure we had already seen or expected:
- Completeness — every string present in every language, so no reader silently falls back to English.
- Script and diacritics — each language must actually show its own characters. This catches stripped accents, which are not cosmetic: in a separate keyword measurement, a Turkish phrase written without its diacritics returned a hundredth of the search volume of the correctly accented form.
- Forbidden wordings — a per-language list of claims we must never make, such as calling an unconfirmed candidate a verified supplier, or calling an outage a company that does not exist.
- Length — a budget per string so nothing overflows a card on a narrow screen.
We verified those gates were not decorative: stripping the Turkish diacritics made the script check fail, and injecting a forbidden phrase made the wording check fail. They work.
And they caught none of the eleven. Every gate here operates on characters, on word lists, or on counts. An error like work permit where business licence belongs is well-formed at every one of those levels. Meaning is simply not visible from that layer, and no amount of stricter character rules would have found it.
This is the finding with the widest application: if your localisation quality assurance consists of automated string checks, you are testing whether the text is intact, not whether it is right.
Method, and what this study cannot say
One language model produced all translations. A second, independent model was then given each string with the intended English meaning and asked to back-translate literally, and to flag only three things: a wrong domain term, a wrong or dangerously softened legal meaning, or an error a native speaker would notice immediately. It was not shown the first model's reasoning. Flags were then classified by hand and either corrected or accepted with the drift written down.
Using the same model to check its own work is close to useless — it reads its own error as the intended meaning. That is the entire reason the reviewer had to be a different model.
Of the eleven, eight were corrected and the live interface carries the corrected wording. Three were accepted with the drift recorded: a low-risk status word where the reviewer's own preferred phrasing was itself unverified, and where further edits in languages we do not read would have added new unchecked risk rather than removing it.
What it cannot say
- It is not a benchmark. 98 short, domain-specific interface strings, one translating model, one reviewing model, one review date. It measures a workflow, not machine translation as a field.
- Neither model is a native speaker. The reviewer flagged what it could see; it also stated plainly that it could not find authoritative sources for some phrases it was judging. Its flags are findings to investigate, not verdicts.
- The true error rate is a floor, not a ceiling. 11.2% is what a second model caught. Errors both models share — a term they are both confidently wrong about — are invisible to this method by construction. A native-speaker review would plausibly find more.
- The per-language split proves nothing. Six languages with flags and eight without, at one to three flags each, is not evidence about those languages.
- We publish this about our own text. These were our strings and our errors, found before shipping. The corrected wording is live; per our editorial policy the localised interface is marked as machine-assisted and has not been reviewed by native speakers.
Citing the machine-translation review
You may quote or reproduce these figures, including commercially, provided the review date, the frame and the stated limits travel with them. Machine-readable metrics with numerators and denominators are in the result file.
What these studies are, and are not
The separate source-availability method page documents paired vantage points, user agents, same-session controls and correction rules; it is a method, not a ninth Dataset. The verification-cost comparison is a dated buyer guide, not a Dataset.
- Frames are stated, not implied. The 264-company frame is three official provincial excerpts of one national batch — elite, state-vetted manufacturers, not a random sample of Chinese suppliers. Rates measured on this frame are conservative bounds for claims like “even top manufacturers change names”; they are not population estimates.
- Registration data was read through licensed commercial data platforms that republish filings originating in the National Enterprise Credit Information Publicity System; person-name fields were discarded at collection. An absence on a platform is not proof of absence in the official record.
- Availability observations are dated, single-connection facts. A status code from August 2026 says nothing about a portal today, and our measurements support no conclusion about reachability from outside China — we have no trustworthy overseas observation point.
Citing this research
You are welcome to quote or reproduce these figures, including commercially, provided the observation date and the stated limits travel with them. The dates are load-bearing: citing a rate or a status code without its date misrepresents it. Link the study page rather than this index, so readers reach the method and limits. Each study page carries machine-readable Dataset metadata; if you need a clarification about method for a piece you are writing, the contact routes are here.
These pages report measurements, not legal conclusions, audits or safety findings, and none of them evaluates any specific supplier a reader is dealing with. For a specific supplier: the free in-browser screen or a dated China-side record check.