Consumer / Food & Beverage
Wine Data Acquisition & Normalisation
1.5 million multilingual wine reviews and hundreds of thousands of sommelier notes, crawled and normalised into structured taste data — the dataset two startups needed to build recommendation apps on.
The problem
North American startups building wine recommendation and make-your-own-wine apps needed something that did not exist as a dataset: real-world data linking wine characteristics to winemaking processes, and real-world data on what actually pairs with what.
The information was out there, scattered across review sites and sommelier notes in multiple languages, written as prose by people describing taste. None of it was structured, and taste vocabulary does not normalise itself.
What I built
Acquisition at scale. Crawled and scraped over 1.5 million wine reviews from across the world, multilingual, plus hundreds of thousands of sommelier notes and professional reviews.
Extraction. Custom NLP models for information extraction and aspect-based sentiment analysis, pulling structured taste and quality attributes out of free-text tasting notes.
Normalisation. The hard part — reconciling how thousands of writers in multiple languages describe the same characteristic into a consistent schema that an application could query.
Analysis. Identified correlations between a wine’s taste profile and the winemaking processes behind it, and between wines and their pairings.
The outcome
- The clients shipped their mobile applications on the back of the dataset.
- The correlation work gave them the substance behind their recommendations — the reason a suggestion is a suggestion, rather than a lookup table.