Data Engineering & Analytics

One set of numbers, and everyone agrees on them.

Most companies this size do not have a data problem. They have a reconciliation problem: four systems that each hold a defensible version of the truth, a month-end that runs on spreadsheets, and one person who knows why the finance number never matches the operations number.

not in scope · regulated lending decisions · replacing source systems · Scoring and decisioning

Data pipeline — source to report Source systems feed a landing zone, which loads into a modelled warehouse, which serves a scheduled report. Data-quality tests run as part of every load, and a failed load alerts rather than reaching the report. Sourcesystems Landing(raw) Model(warehouse) Report(BI report) scheduled · monitored · tests run inside every load
Source systems feed a landing zone, which loads into a modelled warehouse, which serves a scheduled report. A failed load alerts — it never reaches the report.

last reviewed ·

We build the layer that ends the argument — the warehouse, the pipelines that feed it, and the reporting on top — so that a number in a board pack, a number in a dashboard and a number in the source system are the same number, and you can show where it came from.

What this silo does not cover, stated in full rather than left to the rail: we do not build models that make or materially influence a regulated credit, lending or insurance-pricing decision, and we do not replace the source systems that produce your data — we model and pipeline what those systems produce. See Scoring and decisioning.

What this usually looks like before we start

Reporting nobody trusts is worse than no reporting, because you pay for it and route around it.

What changes

The close gets shorter and less manual.
Data arrives on a schedule, reconciles automatically, and fails loudly when it does not. Your team spends month-end reviewing exceptions rather than rebuilding the same workbook.
Every number has a traceable lineage.
For any figure on any report, you can follow it back through the model to the source system and the load that produced it — what makes reporting defensible to an auditor, a lender, an acquirer or a board that has been surprised before.
Definitions live in one place.
"Active customer", "net revenue", "gross margin" are defined once, in the model, and every report inherits that definition. Changing a definition becomes one reviewed change rather than a search through forty workbooks.
Your analysts stop being a bottleneck.
A modelled, documented semantic layer means the people who know the business can answer their own questions instead of queuing for someone who can write SQL.

Scorecards and KPI reporting

last reviewed ·

A hundred dashboards nobody opens is not a KPI programme — it is the same reporting problem this page already describes, wearing a more colourful skin. Scorecards and KPI reporting are part of this work: agreeing the definitions, building the measures once, and making them reconcile to the systems they came from. Inventory and product-level KPIs specifically — Johan Madrid built those as a markets analyst and BI developer at a global consumer-goods company, for Central America and the Caribbean.

company delivery history · founder-asserted ·

past employment · Johan Madrid, named with consent ·

The scorecard and KPI pipeline itself — from source system to a published measure — is set out in full on Tableau and Power BI.

Machine learning

Machine learning is part of this work.

company delivery history · founder-asserted ·

What we build with

last reviewed ·

administration and migration detail · Tableau and Power BI

We work in the Microsoft data stack, because that is where this client base already is and because it keeps the estate on one identity, one security model and one bill:

We run Tableau Server and Tableau Online — administration, migration, and the semantic layer underneath the dashboards — alongside Power BI and SQL Server, not instead of them.

company delivery history · founder-asserted ·

For companies moving off ageing SQL Server infrastructure, this work is frequently the second half of a migration project — see Azure & Microsoft Migration. Migrating the database and modelling the warehouse are adjacent purchases, and doing them in the wrong order wastes money: rehosting a database nobody trusts just relocates the problem.

In finance

Reporting that has to be defensible, not just informative.

  • Management reporting and board packs built from a single reconciled source, with a lineage you can show.
  • Portfolio, exposure and segment reporting with history preserved as it was recorded.
  • Regulatory and client reporting where the number must reconcile exactly, every time.
  • Row-level security so people see the accounts, branches or portfolios they are entitled to see.

In retail

Reporting that has to arrive in time to act on.

  • Sales, margin and inventory across stores, channels and categories, in one model rather than three that disagree.
  • Point of sale, e-commerce and ERP data unified — usually the hardest and most valuable part.
  • Stock, replenishment and shrink reporting current enough to act on.
  • Seasonality and peak analysis using history you can trust.

Governance is part of the build

last reviewed ·

Governance is not a separate phase we bill for later — it is built into the pipelines from the start. Every load runs data-quality checks before it publishes, and cross-system integrity tests catch the same customer, account or SKU disagreeing across sources before a report ever shows it. Every number carries documented lineage back to the source system and the load that produced it. Reference data — customers, products, accounts — is stewarded as master data from one place, so "the same customer" means the same customer in every system that touches it.

company delivery history · founder-asserted ·

What this silo covers

How we work

Every engagement is scoped and quoted individually, with you, before any work starts.

  1. Understand the questions. We start with the decisions the reporting is supposed to support, and trace where today’s numbers come from and who adjusts them by hand.

  2. Assess the sources. What systems hold the data, what condition it's in, and where the quality problems are — including the parts that need cleaning up at the source.

  3. Model it. A warehouse designed for the questions you actually ask, with the grain, conformed dimensions and history you need.

  4. Build the pipelines. Scheduled, idempotent loads that alert on failure and never half-write or duplicate — orchestrated and monitored, with failure recovery designed in: Prefect or scheduled SQL Server jobs, Databricks where volume calls for it.

  5. Report on it. Power BI or Tableau — usually whichever your people already know — plus a semantic model your team can extend.

  6. Hand it over. Documentation, source control, a deployment pipeline, and working sessions with the people who will own it.

company delivery history · founder-asserted ·

Built for

  • Companies with 15 to 500 employees where month-end or reporting depends on manual reconciliation — the industry matters less than the number of systems that have to agree.
  • Teams with real analysts already in place who need an engineering layer underneath them, not a replacement for them.
  • A company running across a handful of source systems, where the problem is modelling and pipelines rather than platform choice.
  • Companies already on, or moving to, the Microsoft data stack.

Not a fit

  • You need a regulated credit, lending or insurance-pricing model — see Scoring and decisioning for exactly where that line sits.
  • You want a report count or a dashboard count delivered — a hundred dashboards nobody opens is not a successful project, and we will say so.
  • Your data volume or streaming requirements genuinely need a lakehouse platform on day one — the assessment will typically tell you if that's true, and often it isn't.
  • You want the visualisation layer fixed without touching what's underneath it — a BI tool is a window, and a new window doesn't change what's behind the wall.

Access to your data during this work follows the same standard we hold ourselves to everywhere else — named individuals, least privilege, logged and time-bound elevation.

Questions buyers actually ask

We already bought Power BI and it did not fix anything. Why would this be different?

Because a BI tool is a window, and the problem is usually behind the wall.

If four systems disagree, if there is no shared definition of "customer", and if the numbers require manual adjustment before anyone believes them, then adding a visualisation layer produces prettier versions of the same disagreement. That is the single most common state we find, and it is why the tool gets blamed.

The work that changes things is the modelling and the pipelines underneath: one reconciled source, definitions agreed once, loads that test themselves, lineage you can trace. Once that exists, the reporting tool you already own usually turns out to be fine.

Do we need Microsoft Fabric, or a data lake, or a warehouse at all?

Probably not all three, and quite possibly none of the fashionable ones.

A company with a few million rows and half a dozen source systems is well served by a properly-modelled Azure SQL database and a scheduled pipeline. That is unglamorous and it works, it costs a fraction of a lakehouse platform, and your team can operate it without specialist skills you would have to hire.

We recommend the heavier platforms when there is a reason — data volume that genuinely does not fit, semi-structured or streaming sources, or a real-time requirement. Not because it is on a roadmap slide. The assessment typically tells you which situation you are in, including when you need less than you think.

How long before we see anything useful?

We aim for a first genuinely useful report early, on a narrow slice — one subject area, end to end, in production, being used — rather than a long build ending in a single reveal.

There are two reasons, and only one of them is about morale. The first is that a working slice is the only reliable way to discover what the data is really like; source system documentation is optimistic almost everywhere. The second is that it gives you a real decision point early, at low cost, about whether to continue.

Who owns the models and the code?

You do, without qualification.

The warehouse is in your Azure subscription. The pipeline and transformation code is in your repository. The Power BI or Tableau content is in your tenant, under your licences. Documentation is written for your team's use, not as a lock-in device.

We do not build things only we can operate. If we vanished, a competent data engineer could read the repository and the documentation and carry on.

Our data is a mess. Should we clean it up before we call you?

No. Assessing the mess is part of the work, and most attempts to tidy data before a project are wasted because nobody yet knows which parts matter.

What does help before a first call: a list of your source systems, a rough idea of volumes, the reports your business actually relies on today, and honesty about which numbers currently get adjusted by hand and by whom. That last one is the most useful thing you can bring, and it is the one people are most reluctant to say out loud.

Can you work with our existing analysts?

Yes, and it is usually the better outcome.

Your analysts know the business, which is the part that cannot be outsourced. What is often missing is the engineering layer underneath — modelling, pipelines, testing, source control, deployment. We build that, and work alongside your people so they can use and extend it.

Where a team wants to learn the practice rather than just inherit the output, we will work that way deliberately: pairing, review, and documentation aimed at handover from the start.

Start with a call about what you are trying to fix.

Book an assessment call

30 minutes · with the engineer who would scope the work

one-page recap within 2 business days of a call that progresses