**Big Data Must Talk Structuring [Dr. Cheong's Talk Part 4]**

**Big Data Must Talk Structuring [Dr. Cheong's Talk Part 4]**

2017-07-11Insight

Image

Before the term "big data" became a buzzword, various types of data were already ubiquitous. Common examples include government-released population and economic data, such as birth rates, unemployment rates, and GDP. These are collected through registration and surveys, then processed and calculated. Additionally, industries generate customer, product, and transaction data, such as the diverse product information, prices, transaction quantities, and amounts on e-commerce platforms, or call durations in telecommunications and credit card transaction details. These are behavioral data resulting from various activities.

With the advent of the internet era, every online action, such as a mouse click or a keystroke, is recorded in the server logs of the websites visited, forming massive and continuously accumulating data. When people borrow books from libraries or check in luggage at airports, their location and presence are tracked by RFID, GPS, or CCTV cameras, generating large amounts of data. If such data reflect people's behaviors, then every like, sigh, photo upload, or ephemeral video posted on social media represents people's emotions and states of mind.

The data mentioned above, from processed data to social data, can be described in big data terms as ranging from "structured" to "unstructured" data.

Structured Data vs. Unstructured Data

"Structured" data refers to data that can be neatly and orderly arranged. A simple example is an Excel spreadsheet, where each entry has fixed columns, formats, orders, and sometimes lengths. For instance, an employee data sheet with hundreds of employees' IDs, names, genders, birthdates, and salaries is recorded in cells formed by intersecting rows and columns. This data can be represented numerically and subjected to arithmetic calculations.

In contrast, "unstructured" data lacks a defined structure. For example, a Facebook post praising a dish, accompanied by mouth-watering photos, has no columns, fixed formats, or immediate numeric representation for arithmetic calculations.

Why Structure Matters

Understanding the structure of data is crucial for making big data useful and valuable. The first step is organizing chaotic data into orderly data, allowing for subsequent processing and analysis based on specific rules or algorithms. Traditionally, many data sets were designed with predefined formats, and collected data were filled in accordingly, such as reports and dashboards generated by business intelligence (BI) systems.

In the big data era, however, data from web usage logs, sensor-collected location and image data, and social media text, images, and videos emerge unpredictably and in real-time. These must be collected and then organized into a structured format to truly activate the data.

Structuring big data lays the foundation for creating value from the data.

Dr. Angus Cheong Chairman of the Asia-Pacific Internet Research Alliance and Chief Data Consultant at uMax Data Technology Ltd.

(Originally published in Hong Kong Economic Times, reprinted with permission)

New

Product

Cases

About

News

English繁體中文
Logout