Back
Foundation of the three demos

Architecture

I take projects all the way to production: not just models, a whole architecture

What runs today for the three demos, and the target architecture a house would run in production: delivery from GitHub, data, models, security and monitoring. Hover a brick to see its links. Figures are measured on the demos' synthetic data.

The target design a house would run: what I would put in place, and what exists today.

01SOURCE02BUILD & RELEASE03RUN04DATA & AI05SECURE06OBSERVEGitHubOne repository, protected mainGitHub ActionsContinuous integrationKeyless accessWorkload Identity FederationArtifact RegistryImmutable imagesTerraformInfrastructure as code, perenvironmentProgressive releaseCloud Run revisionsLoad balancerCloud Armor in frontIdentity & SSOIdentity-Aware ProxyAPICloud Run, privateScheduled jobsCapture and watchBigQueryBronze, silver, goldMLOps pipelineAgent Platform PipelinesGeminiAgent Platform onlyDataiku & BIWhere the business worksLeast privilegeIAMPerimeter & keysVPC Service Controls, CMEKSecretsSecret ManagerAudit & privacyAudit logs, DLPSLOs & alertingCloud MonitoringModel monitoringDrift and live errorFinOpsBudgets and ceilingsLogs & tracesLog Analytics, Trace
01 · Source
02 · Build & release
03 · Run
04 · Data & AI
05 · Secure
06 · Observe
Cloud Run, private

API

The same FastAPI service, one instance kept warm so no visitor waits for a cold start, with one dedicated service account and no public address.

Talks toIdentity & SSO · Progressive release · Terraform · BigQuery · Gemini · SLOs & alerting · Logs & traces · Secrets

Region
europe-west9, Paris
Warm
Minimum one instance
Today
us-central1, scales to zero, two instances at most
Models built

The predictive and decision models behind the demos, apart from the Gemini calls. Each one is measured on weeks it never saw, against the business rule it replaces.

Log-linear regression · BigQuery ML

Demand forecasting

10% vs 21%error per bay-week, model against the register's average

  • sell-out
  • brand effect vs location effect
  • traffic
  • seasonality
  • launch lift
Data
Two years of weekly sales per bay, 40 stores, trained on 78 weeks.
Measured
Wrong by 10% on a bay's weekly sales, against 21% for the 13-week average a merchandiser uses today. In money: on a bay that sells €10,000 a month, the forecast is off by about €1,000 instead of €2,100, so a rent or a reallocation rests on a figure twice as reliable. Where it matters most, a bay that changes brand: 10% instead of 39%. Scored on 26 weeks the model never saw.
Used by
Prices every Store Engine simulation, ranks brands for vacant bays in the reorganisation assistant, draws the forecast curve in each bay's sheet. Readable: Christmas ×2.75 and entrance ×1.83 are recovered from the weights.
Tried and rejected
Boosted trees: 12.4%, worse and unreadable. Adding each bay's own identity: 9.1% overall but 12.2% on new placements instead of 10.3%, because it absorbs the brand effect. Rejected: the simulations are about changing brands.
Statistical model, shrunk to the mean

Event & promotion uplift

17 vs 27 ptsmean error on 58 past events, model against ignoring events

  • real events
  • launch, pop-up, show
  • Paris vs France
  • uplift
  • scarcity
Data
Measured effect of 79 real events (houses and shopping days) on brand bays against their category, by event type and impact.
Measured
Sizing a limited edition with the events instead of last weeks' sales brings the volume 3.6 points closer to real demand at the retailer (12.9% off instead of 16.5%) and 9 points closer in a house's own boutiques (7.6% instead of 16.6%). In money: on a run of 10,000 sets, about 360 fewer sets made in the wrong place at the retailer, 900 fewer in the boutiques. The lift of a coming event is estimated to within 17 points, against 27 when events are ignored, and the stated range holds 3 times out of 4.
Used by
Feeds the allocation: the lift and its spread for each upcoming event.
Tried and rejected
First version of the spread: shrunk with the number of past events, so the 80% band held only 25 of 53 events (47%). The spread of a new event does not shrink with the data: corrected, 79% on the 53 events of the time, 76% on today's 58.
Used in
Allocation
Response curve, least squares

Media saturation curve

0.44 vs 0.96 ptsmean error on 3 held-out months, curve against past average

  • retail media
  • share of voice
  • diminishing returns
  • saturation point
  • marginal return
Data
A synthetic year of monthly media spend per brand and site, 24 brands, 7 sites, each buying share of voice along a hidden curve.
Measured
Predicts the share of voice a monthly budget buys to within half a point, twice as close as a brand's past average (0.44 against 0.96 points, on 3 months it never saw). In money: it tells a media director where the next €10,000 stops paying, and 5 brands of 24 already spend past that point.
Used by
Gives each brand its saturation point and what €10,000 more a month would buy; 5 brands of 24 spend past it.
Generative AI

Where Gemini is used, and why that model. The rule everywhere: Gemini reads, a model measures, the code decides. Every figure on screen comes from a tool or a model, never from what Gemini writes, and every call goes through one gate: a daily cap of $1, a free-tier fallback, one JSON log per call.

Classification · structured output

Gemini 3.5 Flash-Lite

  • article
  • event kind
  • place
  • impact in code
Context
Reads each article about a house: kind of event, place, whether it reaches the retailer's customers. The impact (every store, Paris, none) is then derived in code. Used on the 79 catalogue events and on the 13 Louis Vuitton events, and on every article of the daily watch.
Why this model
A short text in, a few fields out: the cheapest model is enough, with low reasoning.
Measured
Impact placed right on 73 of 79 events (92%) against hand labels, 12 of 13 at Louis Vuitton.
Cost
About $0.02 for the 79 events.
Used in
Allocation
Grounded search

Gemini 3.8 Flash + Google Search

  • news watch
  • sources
  • daily job
Context
Every morning at 7:00 a Cloud Run job searches the news of each of the 7 houses with Grounding with Google Search, drops any article without a source, then hands the rest to the classifier. Nothing is stored but the link and the fields.
Why this model
Flash-Lite answered without returning its sources on the test day, and an unsourced article is not kept.
Measured
A week of the watch is still to be followed: articles kept per house, classification errors, real cost.
Cost
About $0.013 per search, $3 a month.
Used in
Allocation
Agent with tools

Gemini 3.5 Flash-Lite

  • reorganisation
  • tool calls
  • forecast model
  • plan
Context
Reads a store, asks the forecasting model what each brand would sell on the bays worth changing, and proposes a reorganisation. It only chooses: the code rechecks every move (unknown bays, running contracts, wrong department) and reprices the plan.
Why this model
The tools already do the arithmetic. 3.8 Flash answered in 30 to 60 seconds on test days, Flash-Lite in 10 to 16.
Measured
Plans checked and repriced by the same simulation as a visitor's scenario; tested in test_advisor.py.
Cost
$0.003 to $0.006 per plan, measured in Paris and Lyon.
Agent with tools

Gemini 3.5 Flash-Lite

  • questions in words
  • French and English
  • read-only
  • brand isolation
Context
Answers questions about the register in French or English with two read-only tools. Every figure comes from a tool, the rows it read are returned under the answer, and the tools only see what the visitor's profile may see.
Why this model
Two or three short calls with tools: speed and cost matter more than depth.
Measured
Isolation tested: a brand never receives another brand's rows. Ten real questions, three of them out of scope, are still to run.
Cost
About $0.003 per question, estimated, not yet measured.
Vision · structured output

Gemini 3.8 Flash

  • floor plan
  • image import
  • paid tier only
Context
Reads the image of a floor plan exported from a planogram tool and returns walls, entrance and each fixture as boxes. The code converts them to metres from the store's real width, and the editor lets the user correct the draft.
Why this model
Spatial reading needs the stronger model. Paid tier only: the free tier may reuse images, and a visitor's plan must never go there.
Measured
95% of fixtures right (type, centre within 1 m) on 9 drawn plans, 2 to 5 cm position error, department right 71 to 88% of the time on labelled plans.
Cost
About $0.008 and 10 to 20 seconds per plan. 10 imports per IP a day.
Vision · structured output

Gemini 3.5 Flash-Lite

  • retail media
  • sponsored placements
  • brand detection
Context
One call per screen of a retailer page, captured by Playwright: for each sponsored placement the brand, product, format, label and box. The brand is mapped to the universe's brand ids.
Why this model
Stylised luxury visuals defeat classic OCR; a vision model reads the visual in one pass. Flash-Lite is enough at this volume.
Measured
A sample of 56 hand-checked placements: 54 brands right, 55 boxes right, no false placement.
Cost
About $0.0015 per screen.
Vision · set aside

Gemini 3.8 Flash

  • control photo
  • shelf audit
  • facings
  • PLV
Context
The first Store Engine checked shelves from a photo: count of facings, display material against the contract. Set aside on 5 October 2026 when the product became the space register; code kept under the git tag store-engine-v1-execution.
Why this model
Tested on 12 public photos across store, wall and shelf level.
Measured
Facings within two on 88% of shelves, display material found 5 times out of 5, but empty slots 0 of 6. A stronger model finds some, with false alarms at the same confidence: not a basis for an audit.
Cost
About $0.0015 per photo with low reasoning.
What it shows about the way I work
  • FramingEvery demo starts from a question a director asks weekly. The business sets what a lost sale and a set left over cost, from the price and the margin of the product, and every plan is priced with them.
  • Delivery and measureEach model is compared with the business rule it replaces, on weeks it never saw: 10% error against 21%. Terraform, jobs and logs are the hand-over to data engineers.
  • Adoption and alertsBrands and the retailer get their own view, questions are asked in plain words, and spend alerts fire at 25 to 100% of budget.