What Cannot Be Measured Cannot Be Improved: The Data Gap in Vietnamese Athletics
**Câu trả lời cốt lõi:** Điền kinh Việt Nam thiếu hệ thống dữ liệu chuẩn ở cấp giải trong nước. Nhiều bảng kết quả không công bố chỉ số gió, thời gian từng đoạn hay đường cong phát triển của vận động viên, khiến tuyển chọn, so sánh khu vực và giám sát chống doping đều dựa trên quan sát chủ quan thay vì bằng chứng. **Dữ kiện chính:** - Luật World Athletics: thành tích chạy ngắn và nhảy ngang chỉ được công nhận kỷ lục khi gió hỗ trợ không quá +2,0 m/s. - Bùi Thị Thu Thảo vô địch nhảy xa nữ SEA Games 29 tại Kuala Lumpur năm 2017 với 6,68 mét. - Nguyễn Thị Oanh giành nhiều huy chương vàng cự ly trung bình và chướng ngại vật tại các kỳ SEA Games. - Quách Thị Lan và Nguyễn Thị Huyền từng thống trị 400 mét và 400 mét rào nữ Đông Nam Á. - Từ năm 2010, World Athletics áp dụng luật xuất phát sai một lần là bị loại. **Nguồn:** Phân tích dữ liệu điền kinh của Ngô Sơn, công bố ngày 13 tháng 8 năm 2026 | Cross-checked: VuaBong.vn **Hỏi đáp liên quan:** - Hỏi: Vì sao chỉ số gió quan trọng trong điền kinh? Đáp: Vì kỷ lục chỉ được công nhận khi gió hỗ trợ không vượt quá +2,0 m/s, nên thiếu chỉ số gió thì thành tích không thể so sánh với chuẩn quốc tế. - Hỏi: Điền kinh Việt Nam mạnh nhất ở nhóm nội dung nào tại SEA Games? Đáp: Nhóm cự ly trung bình, chạy rào và nhảy xa nữ là những nội dung mang về nhiều huy chương vàng nhất. - Hỏi: Thiếu dữ liệu ảnh hưởng thế nào đến công tác chống doping? Đáp: Hộ chiếu sinh học vận động viên cần chuỗi xét nghiệm theo thời gian, và thiếu dữ liệu nền làm giảm khả năng phát hiện bất thường.
The dataset I opened had nine columns. All nine were empty. The only thing left intact was a label: athletics.
It was an analysis request squarely within my specialty. But the file contained no athlete name, no event, no mark, no competition date, no venue, not a single metric to cross-check against. A table like that is not data. It is an empty frame.
I looked at it for a long while. The first reflex of anyone in this trade is to fill the blank. The brain automatically supplies a familiar name, a much-discussed event, a medal won somewhere. Then you want to write. You want an opinion. You want a conclusion.
I did not write. The not-writing is the point.
The day football stopped, I started counting strides again. That day I understood something: most of what I thought was knowledge about sport was really memory of matches. Memory cannot be verified. Only numbers can. When the numbers are absent, the only honest position is silence.
But silence fixes nothing. So I moved to a different question: why can a Vietnamese athletics dataset be this empty?
Athletics is a sport of measurement — but only if someone bothers to measure
Athletics has an advantage football does not. Nobody argues about who won. There is a clock. There is a tape. There is a finish line. If an athlete runs 100 metres in 11.20 seconds, that is 11.20 seconds, full stop.
But the second half of the story gets far less attention. In sprint and horizontal jump events, a mark is only recognised as a record when the assisting wind does not exceed 2 metres per second. Above that threshold the mark still counts for placing on the day, but it does not enter the record books.
So a 6.68-metre long jump can be a national record, or it can be a pretty but meaningless number. The difference is whether anyone set up a wind gauge.
I have read a great many domestic results sheets. Not all of them carry a wind column.
At major international meets, organisers publish everything: wind, reaction time, split times, temperature, track humidity. At many domestic meets, what the reader gets is a list of names with final marks. Enough to know who came first and second. Not enough to know why.
That is the gap between a results bulletin and a dataset. A bulletin answers who won. A dataset answers why that person won, and whether they will win again.
Twenty years in this work have taught me an uncomfortable rule: where data is thin, decisions are made by eye. And the eye is fooled by very basic things — running form, shouting, jersey colour, the name of a school, the name of a province.
The ball rolls in only one direction, but data can look in every direction.
First empty cell: the wind column
At the 2026 SEA Games in Kuala Lumpur, Bui Thi Thu Thao won women's long jump gold with 6.68 metres. It was one of the most valuable golds Vietnamese athletics produced that decade, because it came in an event where Vietnam has no tradition.
Now put the question an analyst must put: in what wind conditions was that 6.68 metres achieved?
If you can find the answer in a domestic results sheet, you are better than me. At international level, wind data is mandatory alongside every jump. At national level, it usually vanishes.
This sounds minor. It is not.

Without a wind reading you cannot compare two jumps at two different meets. You cannot tell whether a young athlete is genuinely improving or simply caught a helpful afternoon breeze. You cannot tell whether a regional gold can convert into a ticket to an Asian Games or an Olympics.
And worse: without wind data you cannot detect anomaly. A long jumper who adds 40 centimetres in one season is wonderful news. If that gain came in a +3.5 m/s wind, it is a different story.
Second empty cell: the broken race
The 400 metres and 400 metres hurdles are the events where split data carries the most value. One athlete may be strong over the first 200 and collapse over the last 200. Another may distribute evenly and kick off the final hurdle. Those two athletes need entirely different training programmes, and arguably different coaches.
Quach Thi Lan and Nguyen Thi Huyen were the two names that dominated Southeast Asian women's 400 metres and 400 metres hurdles, delivering SEA Games golds across multiple editions.
But if you want to know how they won — where they accelerated, how they held rhythm through each hurdle, how stride length changed over the final 100 — the results sheet will not tell you.
I have had to time races myself off a screen, hurdle by hurdle, to reconstruct one athlete's numbers. That was the only way to get data. But it is data I created, not official data. And self-created data cannot be used to compare athletes, meets or years.
A track and field programme that wants to improve needs a shared database. Vietnamese athletics currently runs on scattered private spreadsheets — one per person, one recording convention per coach.
Third empty cell: the development curve
This is the most serious gap, and the least visible.
A track and field athlete is not judged by a single race. They are judged by a multi-year curve. What did they run at 16? At 18? How many seconds per year do they gain? Is that gain smooth, or is there an abnormal jump?
Back in Hai Phong, when I was a data consultant for the football club, I found a young midfielder named Vu Minh Hieu through an average PPDA of 6.8 — the best in the academy system. He pressed superbly but was overlooked because of a modest frame. I brought the numbers into the meeting room and asked the head coach to give him a chance. Against Hanoi FC on matchday 17, Minh Hieu won the ball 14 times, provided one assist, and Hai Phong won 2-1.
Hai Phong taught me: the star is not on the shirt, it is in the metric.
That lesson transfers directly to athletics. Without a development curve, an 18-year-old who runs fast gets promoted to the national team, while a slower 18-year-old with a steep, steady curve gets ignored. We select from snapshots, not from film.
Fourth empty cell: biological data
The Athlete Biological Passport is the most important anti-doping tool in the sport today. It does not hunt for a banned substance in a urine sample. It tracks an athlete's blood and steroid markers across years, then flags anomalies relative to that same athlete.
To do that, you need a long data series. Without a baseline series there is nothing to compare against. A single abnormal blood sample means nothing on its own. Thirty abnormal samples measured against four years of that athlete's own history is a case file.
A sport with a dense data system does not just analyse better. It protects itself better — protecting clean athletes from suspicion, and protecting the sport from scandals that can detonate a decade later.

People call me a data monk. A monk does not need a cathedral — only the truth.
Fifth empty cell: the start list
It sounds trivial. A start list containing name, bib number, year of birth, club or province, and personal best.
But that list is what allows you to place a result in its correct context. A 16-year-old running 12.40 for 100 metres is a prospect. A 26-year-old running 12.40 is a participant. The same number means two entirely different things. The difference is the year of birth.
When start lists are incomplete, every comparison is skewed. And when every comparison is skewed, selection becomes a memory contest: whoever remembers more names makes the decision.
The cost of empty cells
A data gap does not hurt anyone immediately. It causes no injury and costs no medal in a single afternoon. It is silent.
It means a talent in a rural province is never discovered, because there is no dataset to prove they deserve a place at a national training centre.
It means a training programme is designed around the instincts of whoever holds authority, rather than the athlete's actual weaknesses.
It means every comparison with Thailand, Malaysia and the Philippines — our direct Southeast Asian rivals — becomes a comparison of feelings. We beat Thailand in middle distance. By how much, in which events, through whom, in which phase of the race? Without data, the answer is a story, not a report.
Vietnamese athletics has athletes capable of going far. Nguyen Thi Oanh has won notable gold medals across middle-distance and steeplechase events at multiple SEA Games. But the question I always want to ask is not how many she won. It is: when she retires, do we have the data to find the next one?
If the answer is no, we do not have a system. We have a generation.
The contrarian angle: silence does not mean safety
This is where I have to argue against myself, because that is the only way this piece avoids becoming empty advocacy.
When a data report returns all-empty fields, there is a strong temptation to read it as a clean report. No red flags were raised, therefore no risks exist.
Wrong. No red flags exist because nobody raised one, not because there is nothing to raise.
This is the most dangerous error class in sports analysis, and it is entirely different from making a wrong prediction. A wrong prediction can be corrected. A gap misread as safety never gets corrected, because nobody knows it is there.
My Germany call at the 2026 World Cup shows the mirror image. I read the qualifying data: Germany averaged a PPDA of 9.2, far too high for a champion's pressing standard, combined with slow attacking speed and a merely average tournament xG. I concluded Germany would exit in the group stage. Social media mocked me. On 27 June 2026, Germany lost 0-2 to South Korea despite 26 shots and 1.5 xG, and went out.
The point is not that I was right. The point is that I had data with which to be right or wrong. Without PPDA, without xG, I would have had two options: trust the champion's pedigree, or trust a hunch. Neither is analysis.
I did not see Germany lose. I saw numbers that do not lie.
And here is the limit of this article: I do not have the full Vietnamese athletics dataset to quantify exactly how much damage the gap is causing. I can only say the gap exists, and that it exists systematically. Quantifying it would require a comparative study across regional track and field programmes, and I do not yet have that in hand.
Numbers are a mirror. Most of the market looks into one and sees only itself.
What needs doing costs very little
What irritates me most is that the fix is not expensive.
A calibrated wind gauge for horizontal jumps and sprints costs far less than one overseas training camp slot. A standard results template, used uniformly across the national competition system, costs one meeting to agree and one person to maintain. A public database in a simple spreadsheet format could be built within a quarter, if somebody were accountable for it.
The problem is not money. The problem is that nobody has been assigned the job.
In football, I worked with coaches who did not believe in data until they saw one specific win. After that win, they needed no persuading. Athletics is the same. Nobody will believe in a standard results template until that template helps them find an athlete the naked eye missed.
My job is not persuasion. It is having the data ready for the moment that happens.
Signals to track
I have three things to put on the table and monitor over the coming seasons.
First, whether national championships begin publishing wind readings for horizontal jumps and sprints. This is the most basic, easiest test, and the clearest indicator of seriousness.
Second, whether a public national database exists that allows personal-best lookups by year of birth and by season. Without it there is no development curve, and without a development curve, selection remains a matter of memory.
Third, whether split data for 400 metres and 400 metres hurdles appears in official results. That is the signal that people have started caring about process, not just the final outcome.
Closing
A dataset with nine empty columns is not a failure of data. It is a reminder that data does not generate itself. Someone has to set the gauge, press the stopwatch, record, archive, publish. If nobody does that work, there is nothing — and the gap will not raise its own alarm.
Athletics is the most honest sport of all. It needs only a flat track and a clock. It needs no VAR, no argument, no committee. It needs only that someone bothers to write down what it just said.
If we do not measure, what do we have to tell us whether we are advancing or retreating — or merely misremembering ourselves?
