Portfolio

Data analysis

Machine learning, exploratory analysis and advanced queries. Python and SQL applied to real problems where the data needs structure, cleaning and a model that answers specific questions.

01

Machine Learning

Predictive and classification models applied to specific business decisions. Each project starts from a specific question and ends with a metric that measures whether the answer works.

Do all your customers deserve the same offer?
The problem
An electronics chain treated all its customers the same way: the same offers, the same emails and similar discounts. With 6,000 customers and 26 variables, it needed to find behaviour patterns that would let it personalise its campaigns.
The solution
I selected 9 behavioural variables related to spend, frequency and channel. The analysis identified 4 customer profiles, from premium to at risk of leaving, and matched a strategy to each one.
The result
96.8%accuracy in the premium segment
K-Meanst-SNE
View →
What happens to sales when something unexpected happens?
The problem
A retail chain could forecast normal weeks fairly well, but its predictions failed when promotions, strikes or logistics problems came up. The history explained the past, but not always what was happening around it.
The solution
I compared two models over 156 weeks: one based only on historical sales and another that added five external variables. I validated them with time-based backtesting and simulated more than ten future scenarios.
The result
−32%error compared with the history-only model
SARIMAXstatsmodels
View →
Who are we about to lose?
The problem
An airline sorted its passengers into three tiers with rigid rules: Basic, Frequent and Premium. But those rules didn’t always reflect how people actually behaved, so it was hard to know who really needed a loyalty action.
The solution
I tested three models on 21 variables covering flights, spend and incidents, validated them against a baseline and analysed which variables helped tell the different passenger profiles apart.
The result
×2improvement over the previous rules
Gradient Boostingscikit-learn
View →
How much will we sell next week?
The problem
A retail chain plans stock and staffing mainly by looking at what it sold the year before. The problem: ordering too much creates excess stock; ordering too little means running out of product when demand rises.
The solution
I trained a forecasting model on 208 weeks of sales and validated it on another 52 weeks the model had never seen. The aim was to test how far the history could anticipate future demand.
The result
2.21%average forecast error (MAPE)
SARIMAstatsmodels
View →
What makes a customer choose a flight?
The problem
An airline combines price, stopovers, luggage and flexibility, but doesn’t know how much each factor weighs in the customer’s decision. And that weight can change with the type of passenger.
The solution
I analysed 24 flight combinations rated by 1,000 customers. Using conjoint analysis, I broke down 24,000 ratings to measure how much each attribute contributes, both for customers as a whole and for each segment.
The result
69%weight of price and stopovers in the decision
Conjointstatsmodels
View →
See more machine learning projects on GitHub
02

Python

Exploratory analysis with Python, applied to specific business questions. Each project takes a dataset, cleans it, analyses it and ends with a finding that changes the conversation.

EDA · Sales
The problem
A services company with three departments needs to decide where to strengthen its team. The assumption is that the longest-serving employees sell more, but nobody has checked it with data.
The solution
Null diagnosis, justified imputation, productivity ratios by department and a correlation matrix to test whether age predicts sales.
The result
-0.1age–sales correlation
PandasNumPySeaborn
View →
EDA · Education
The problem
A university’s guidance service wants to know which habits best predict exam marks, in order to design a tutoring programme aimed at students at risk.
The solution
Categorisation with pd.cut(), group comparison by median and percentiles, and a panel of 6 visualisations: histograms, scatter plots, heatmap and pairplot.
The result
+22points for studying > median
PandasMatplotlibSeaborn
View →
EDA · Health
The problem
A health insurer wants to cut costs by identifying patients at high cardiovascular risk early, and to decide whether it’s fair to adjust premiums by risk profile.
The solution
Converting categorical variables to numeric, ranking risk factors by sorted correlation and segmenting patients with a multi-metric groupby.
The result
-0.94steps–risk correlation
PandasNumPyMatplotlib
View →
See more Python projects on GitHub
03

SQL

Advanced queries, data exploration and analysis with T-SQL. Each project is self-contained: it creates its tables, loads realistic data and solves business problems with CTEs, window functions and subqueries.

Sales analysis · Automotive
The problem
A network of 50 Volkswagen Group dealerships needs to identify its best salespeople, compare performance across brands and segment dealerships by turnover.
The solution
Chained CTEs, correlated subqueries, LAG() for month-on-month change, PERCENTILE_CONT for distribution and RANK() for rankings. Every exercise solved with two approaches.
The result
10exercises · two approaches
T-SQLCTEsSubqueries
View →
Window functions · Manufacturing
The problem
AdventureWorks needs to analyse its customers’ buying behaviour, rank products by price and measure year-on-year growth per customer to identify the top 10.
The solution
ROW_NUMBER, RANK, DENSE_RANK, NTILE, running totals with ROWS BETWEEN, LAG/LEAD for consecutive rows and CTE + LAG for YoY growth. A final exam in the style of a technical interview.
The result
11+progressive exercises
T-SQLWindow FunctionsLAG/LEAD
View →
Customer analysis · Retail
The problem
A grocery shop needs to segment customers, classify products, spot buying patterns and build cross rankings by city, category and supplier.
The solution
LEFT JOIN to find gaps, CASE for business classifications (VIP/standard, frequent/occasional), CTEs + window functions and advanced combinations: top 3 customers per city, an SQL dashboard.
The result
40exercises · 8 blocks
T-SQLCASERANK
View →
See more SQL projects on GitHub