Data
AI
The Stack Underneath


Most companies think about data team structure only after the first few hires are in place. By then, your data engineer, two analysts and data scientist are working off different definitions of the same metrics. Two dashboards show two different revenue numbers in the same leadership meeting. Nobody owns the gap.
That gap is a structure problem. Structure decides who owns pipelines, who defines metrics and who answers when the numbers disagree. Without that clarity, even a strong enterprise data strategy stalls at execution. Senior hires end up spending their time on rework.
This guide covers the core data team roles, the four common structure models and how your structure should change as the company grows. It also covers who to hire first when building a data team, and when outside help makes more sense than another full-time hire.
A data team structure is the way a company organizes its data people, defines their roles and sets their reporting lines. It also decides how the data team works with the business units that depend on it.
In practice, a data team structure answers four questions.
Who builds and maintains the pipelines and the data platform
Who defines business metrics and keeps them consistent across reports
Where analysts and data scientists sit, in one central team or inside business units
Who sets data priorities and owns the outcome when data work misses its goals
The answers change with company size. A 50-person startup with one data engineer has a structure too. Usually an informal one. That works until the second or third hire starts duplicating work. An enterprise with 40 data people needs the answers written down. Every extra hire raises the cost of overlap.
Structure also sits downstream of strategy. Your data strategy framework sets what the business needs from data. The structure decides which people deliver it. Teams that design the structure first often end up with roles that don't map to any business goal.
Most data team roles fall into three groups. Some roles move and store the data. Some turn it into answers. Others govern it and set direction. The nine roles below are the ones that show up most often, with notes on when each one becomes worth a full-time hire. Few companies need all nine. A team of four often covers six of these roles between them.
A data engineer builds and maintains the pipelines that move data from source systems into your warehouse or lakehouse. They own ingestion, transformation jobs, orchestration and pipeline reliability. When a dashboard goes stale at 9 a.m., the data engineer is usually the first person paged.
This is the foundation role. Every analyst and data scientist downstream depends on the data they deliver. A reliable data pipeline development practice starts with this hire.
The exception is a company whose data lives in two or three SaaS tools with managed connectors. In that setup, an analytics engineer can cover ingestion for the first year.
An analytics engineer turns raw warehouse tables into clean, modeled datasets that analysts and BI tools can trust. The role sits between data engineering and analysis. It became common alongside transformation tools like dbt.
An analytics engineer typically owns three things.
The transformation layer and data models in the warehouse
Metric definitions, so "revenue" means the same thing in every report
Data tests and documentation for the models they build
For many mid-sized companies, this is a stronger first data hire than a pure analyst. The trade-off is cost. Analytics engineers are harder to find and usually cost more than analysts.
A data analyst answers business questions with data. Sales wants to know why conversion dipped in March. Marketing wants to know which channel brings in the highest-value customers. The analyst pulls the data, runs the analysis and explains what it means for the decision at hand.
The best analysts spend more time with business teams than with other data people. That's why many companies embed analysts in sales, finance or product once the team grows past a handful of people. In a small team, the analyst often builds dashboards as well. That overlap usually ends once a BI developer joins.
A BI developer builds and maintains the reporting layer the rest of the company uses every day. That includes dashboards in tools like Power BI, Tableau or Looker, the semantic layer underneath them and the access rules on who sees what.
The difference from an analyst comes down to output. An analyst produces answers to specific questions. A BI developer produces reporting assets that hundreds of people rely on without asking anyone. You need a dedicated BI developer once dashboard requests start crowding out analysis work. Before that point, the role is usually part-time work for an analyst or analytics engineer.
A data scientist builds statistical and machine learning models to predict outcomes or explain patterns. Churn prediction, demand forecasting and pricing models are typical examples. The work is experimental. Many models never reach production.
Hiring a data scientist too early is the most common sequencing mistake we see. Without clean, modeled data, a data scientist spends most of the first year doing data engineering work. That's an expensive use of a specialist salary. We've walked into this exact setup on more than one client engagement.
An ML engineer takes models into production and keeps them running reliably. Their work covers deployment pipelines, model serving, monitoring and retraining.
The split from data science starts to matter once more than one or two models are live.
The data scientist decides what the model should predict and proves it works.
The ML engineer makes sure it keeps working at production scale.
If no models are in production yet, you don't need this role. A data scientist with solid engineering skills can ship the first one or two.
A data architect designs how your data platform fits together. That covers storage choices, integration patterns, enterprise-level data models and the standards other engineers build against. The role matters most during a platform migration or when several business units share one platform.
Smaller teams rarely need a dedicated architect. A senior data engineer usually makes these calls. The role earns its place once architecture decisions start affecting several teams at once.
A data steward owns the quality, definitions and access rules for a specific data domain, such as customer or finance data. A governance lead coordinates stewards across domains and sets the policies they enforce. In regulated sectors like fintech and healthtech, this role tends to arrive earlier than elsewhere.
Stewardship is often a part-time responsibility held by someone in the business. A finance controller may steward finance data, for example. Before assigning these roles, it helps to be clear on where data governance ends and data strategy begins. The two get mixed up often.
A head of data owns the data strategy, the team structure and the business outcomes of data work. In larger enterprises, the title is often Chief Data Officer or Chief Data and Analytics Officer. This person sets priorities, manages the budget and answers to leadership when the numbers disagree.
In our experience, most companies bring in a dedicated data leader once the team reaches four to six people. Before that, the CTO or head of engineering usually holds the role informally. A late leadership hire tends to leave each data person reporting to a different department with no shared standards.
Most companies use one of four data team structure models. Each one trades central control against speed inside the business units. The right choice depends on your team size, how many business units consume data and how mature your standards already are.
|
Model |
How it works |
Best for |
Where it breaks |
|
Centralized |
One data team serves the whole company and reports to a single data leader |
Companies with fewer than 10 data people |
Business units wait in a shared queue, and analysts lack domain context |
|
Embedded (decentralized) |
Data people sit inside business units and report to those unit heads |
Companies where each unit has very different data needs |
Metric definitions drift, and the same pipelines get built twice |
|
Hub-and-spoke (hybrid) |
A central team owns the platform and standards, while embedded analysts serve each business unit |
Mid-sized and large companies with three or more data-heavy units |
Embedded analysts end up with two managers unless reporting lines are clear |
|
Federated (center of excellence) |
Domain teams own their data products, and a small central team sets standards and runs the shared platform |
Large enterprises with mature governance |
Falls apart quickly without strong standards and clear data ownership in each domain |
Most teams move through these models in order. A centralized setup works early because one team can set standards fast. Hub-and-spoke usually follows once business units start asking for dedicated support. Federated only works after governance is already mature.
In practice, reporting lines matter as much as the model you pick. An embedded analyst who reports to the head of sales will put sales requests ahead of shared standards. That's reasonable behavior from the analyst's side. It's also where metric drift usually starts. For embedded and hub-and-spoke models, a dotted reporting line to the central data leader keeps definitions consistent. In a centralized team of three, that extra layer adds overhead with no real benefit.
Data team structure usually moves from one generalist reporting into engineering toward a federated model with domain teams and a central platform. Headcount is a rough guide to where you sit. Your position on the data maturity model is a better one, since a 20-person data team with no shared metric definitions still behaves like a much smaller team.
Here is how the structure typically shifts at each stage.
One or two people. The first data hire usually reports to the CTO or head of engineering. There is no formal structure yet. The job is to make the three or four metrics the business runs on reliable and consistent.
Three to six people. This is a small centralized team. Engineering and analysis start to separate, with one person owning pipelines and models while others focus on reporting and analysis. A dedicated data leader usually joins toward the end of this stage.
Seven to fifteen people. The team shifts toward hub-and-spoke. Analysts move closer to business units, either embedded full-time or assigned as dedicated partners. Leads appear for engineering and analytics, and part-time data stewards take ownership of the most critical domains.
Fifteen or more people. Larger enterprises move toward a federated model. A central platform team runs the shared infrastructure and standards. Domain teams in finance, product or operations own their data products. Governance becomes a formal function with its own lead, and the head of data often becomes a C-level role.
Few companies move through these stages cleanly. Most restructure late, after duplicated work or conflicting reports force the change. The better trigger is the work itself. When analysts spend more time waiting on pipelines than analyzing, or two teams build the same model twice, the current structure has run its course.
When building a data team, hire a data engineer or analytics engineer first, then a data analyst, and add a data scientist only after your data is clean and modeled. Hiring out of this order is the most expensive mistake in the process. The specialists you bring in early end up doing foundation work at specialist rates, and the foundation still gets built slowly.
The table below lays out a typical hiring sequence. The "hire when" column matters more than the order itself. It describes the signal that tells you the next hire will be fully used.
|
Order |
Role |
Hire when |
Delay if |
|
1 |
Data engineer or analytics engineer |
Data sits in several systems and reports are assembled by hand in spreadsheets |
Your data lives in two or three SaaS tools with managed connectors, in which case start with an analytics engineer |
|
2 |
Data analyst |
Pipelines are stable and business teams are queuing up questions |
Your first hire still spends most of the week fixing pipelines |
|
3 |
Second engineer |
Pipeline maintenance crowds out new data work |
Managed ingestion tools cover most of your sources |
|
4 |
Head of data |
The team reaches four to six people, or data spend grows faster than results |
The CTO still has bandwidth and the team is under four people |
|
5 |
BI developer |
Dashboard requests start crowding out analysis |
Self-serve reporting already works for most teams |
|
6 |
Data scientist |
Modeled data exists and a specific prediction problem has a clear business case |
Nobody can name the first model the business needs |
|
7 |
ML engineer |
More than one or two models are running in production |
The first model hasn't shipped yet |
|
8 |
Data steward or governance lead |
You handle regulated data, or metric disputes reach leadership |
You have fewer than ten data people and no regulatory pressure |
This order can feel slow to leadership that wants AI results this year. The sequence itself rarely changes. What changes is how fast you move through it, and outside help can compress the first two or three steps from a year into a few months.
One exception is worth naming. If machine learning is your product, as it is for many AI-first startups, an ML engineer may belong among the first two hires. The sequence above fits companies where data supports the business, which covers most enterprises.
For most companies, a blended model works best. A small in-house core owns strategy, business context and metric definitions. Outside specialists handle platform builds, migrations and surge work. The right mix depends on how steady your data workload is and how quickly you need results.
Build in-house when you have enough data work to keep a team busy for the next two years or more. In-house teams build deep business context that no partner can match. The catch is time. Hiring a senior data engineer can take several months, and the first year of a new team often goes into foundations rather than visible results. Building purely in-house is a poor fit if leadership expects outcomes within a quarter.
Outsource when the work is a defined project with a clear end, such as a cloud migration, a pipeline rebuild or a first BI rollout. A partner brings people who have done the same build before, so the timeline is shorter and more predictable. This approach breaks down when nobody in-house owns the outcome. Once the project ends, the knowledge walks out with the partner team.
Blend when you need both speed and long-term ownership. A common setup pairs an in-house head of data or analytics engineer with an outside team that builds the platform and pipelines. The in-house core reviews the design decisions and takes over operations through a planned handover. This is the model we see work most often for mid-sized companies moving from stage two to stage three.
As a data engineering services company, we have an obvious interest in this question. Even so, we advise clients to keep strategy and metric ownership in-house from day one. A partner can build the platform faster. Deciding what "revenue" means for your business has to stay with your own team.
Design your data team structure around your data strategy, so every role maps to a business goal.
Hire a data engineer or analytics engineer first, and add a data scientist only once modeled data exists.
Start with a centralized team, then move to hub-and-spoke when business units need dedicated support.
Restructure based on how the work is flowing, not on headcount alone.
Keep strategy and metric ownership in-house, even when an outside team builds the platform.
If you're planning your first few data hires or your current team has outgrown its structure, start by mapping each role against the outcomes your strategy needs. Gaps and overlaps show up quickly once that map exists. Our enterprise data strategy team works with companies on exactly this step, from defining the target structure to sequencing the hires and platform work that follow.
You might also like