**Big Data Must Talk Quality [Dr. Cheong's Talk Part 3]**

**Big Data Must Talk Quality [Dr. Cheong's Talk Part 3]**

2017-07-10Insight

Image

The value of big data lies in its characteristics of real-time recording, accumulation, calculability, traceability, and reusability. This has led many to form a stereotype about big data: the more data, the better, and if the data volume is large enough, conclusions can be drawn.

Is more data always better? Google.org once launched an online flu prediction platform called Google Flu Trends (GFT). It operated by using algorithms to predict flu outbreaks based on 50 million flu-related keywords searched by users on Google. These predictions were then compared to the known flu incidence reports from the Centers for Disease Control and Prevention (CDC). Researchers found that from 2004 to 2009, GFT's data was remarkably consistent with CDC's data. This led to some claims that algorithms alone could produce predictions consistent with CDC, suggesting that we no longer need to seek the underlying causes of phenomena, as long as there is a statistical correlation.

Too Much Noise Leads to False Correlation

Later, scholars discovered that in 2013, GFT's predicted data was twice that of CDC's reported data. This shift from initial high consistency to later significant discrepancies was thought to be due to the presence of too much noise in the keywords. Many keywords appeared related to the flu but were actually irrelevant, resulting in "false correlations." For example, the search frequency and timing of "high school basketball" and "flu" matched closely, leading to basketball fans being mistakenly identified as flu patients. Possibly for this reason, GFT stopped publishing predictive data online in 2016.

Currently, using keywords for data collection and analysis is the most common practice in the field of text big data. Many public opinion monitoring and brand monitoring charts (commonly word clouds) and reports are based on keywords. Based on years of practical experience, I, Dr. Cheong, must emphasize the importance of careful testing when using "keywords" and establishing a rigorous data cleaning mechanism to ensure subsequent analysis is based on high-quality datasets, free from irrelevant noise data.

To illustrate, during the Chief Executive election, I and my team followed the trend by systematically collecting, cleaning, and analyzing online opinions about the candidates. When using keywords, we adopted a "concept" approach, treating a candidate's name as a concept that includes nicknames or aliases. We also repeatedly tested to exclude noise that could be mistaken for the Hong Kong Chief Executive election or related to a specific candidate. This ensured that the data used for analysis on the big data mining platform was highly representative and relevant.

In conclusion, big data is not about being "big"; data quality is key.

Dr. Angus Cheong Chairman of the Asia-Pacific Internet Research Alliance and Chief Data Consultant at uMax Data Technology Ltd.

(Originally published in Hong Kong Economic Times, reprinted with permission)

New

Product

Cases

About

News

English繁體中文
Logout