# Data Scientists & Data Engineers: How the Best Teams Work

Beverly Wright, Wavicle Data Solutions & Sadie St. Lawrence, Human Machine Collaboration Institute / LinkedIn Learning & Joe Reis, Ternary Data & Victor Cuadros, Microsoft | DE4AI 2024 | 27:55

Source: https://www.youtube.com/watch?v=A9oLe3bqEpY
Channel: MLOps Community, now AAIF Live (https://www.youtube.com/@AAIFLive-x1r). Summarised by MLOps Talks.
Page: https://mlopstalks.com/talks/data-scientists-data-engineers-how-the-best-teams-work
Published: 2024-10-09
Tags: data-engineering, engineering-culture

## TL;DR
- Data science education has focused heavily on models, while many practitioners spend much of their working time on SQL and data pipelines.
- Data science and data engineering teams struggle when they work as separate silos without a shared view of the product they are building.
- Pair programming, learning the other team's work, and meeting in person can improve communication and empathy between the two groups.

## Summary
The panel traces the split between data science and data engineering. Joe Reis describes data engineering as work that originally supported data scientists by building infrastructure and foundations. Sadie St. Lawrence says data science education often emphasized modeling and visualization while giving too little attention to SQL and pipelines. The panelists argue that teams fail when they work in silos, lack a shared picture of the end product, or pass poorly defined work between functions. They recommend pair programming, cross-team experience, and learning enough of the other discipline to understand its needs. AI tools may let one person work across more of the stack and provide faster feedback, but the panel expects the work to change rather than disappear. On remote collaboration, they prefer a mix of remote work and in-person meetings, especially for difficult problems where whiteboards and direct interaction help.

## Key ideas
### Data engineering grew out of the need to support data science
[03:00](https://www.youtube.com/watch?v=A9oLe3bqEpY&t=180s)
Joe Reis describes the recent form of data engineering as emerging partly by accident. Data scientists needed infrastructure and foundations that they could not build or maintain alone, so data engineers took on that work. He says data engineering has since become a distinct and widely used job category. Sadie St. Lawrence adds that early data science roles often combined data engineering, analysis, and modeling in one person. The amount of work involved became too large, so the responsibilities split into separate roles. The panel treats this history as a reason both groups need to understand what the other actually does.

### Data science training often underprepares people for pipelines
[03:53](https://www.youtube.com/watch?v=A9oLe3bqEpY&t=233s)
Sadie says her data science master's program focused heavily on model building and visualization. SQL appeared in the curriculum, but only as a small part of the training. When she entered industry, she found that much of the work involved SQL and pipelines rather than model development. She later taught SQL because she viewed it as a core skill. Joe connects this experience to older curricula that emphasized Python and modeling while treating SQL as less important. Victor says he also encountered a large gap when he moved from software engineering into data science and had to learn what the models and targets meant.

### Separate teams fail when they do not share the product picture
[08:07](https://www.youtube.com/watch?v=A9oLe3bqEpY&t=487s)
The panelists describe siloed teams as a reliable way to create division. Joe compares data engineering and data science groups that only interact at occasional company events. Upstream dependencies are ignored, project and product managers may not align the functions, and no one develops empathy for the work on the other side. Sadie says each contributor may understand their own puzzle piece without seeing the finished puzzle. That makes it hard to know where the work fits or what another team needs. The panel recommends a shared view of the end product, clearer collaboration, and pair programming across the boundaries.

### AI tools may widen individual scope and bring roles closer together
[13:14](https://www.youtube.com/watch?v=A9oLe3bqEpY&t=794s)
The panel considers whether AI tools could lead to more full-stack data professionals. Victor expects the gap between data science and engineering to narrow as both areas become easier to handle. Joe connects this possibility to streaming data, more modular products, interoperability, and shorter feedback loops between software, data, and machine learning. Sadie says AI tools give her immediate feedback when she builds something, instead of waiting days for another team to return a data model or implementation. She expects people's scope to expand rather than their jobs simply disappearing. Victor compares this with the shift from building physical data centers to managing cloud systems.

### Remote work makes communication harder for cross-functional teams
[18:20](https://www.youtube.com/watch?v=A9oLe3bqEpY&t=1100s)
Victor says remote work can interrupt the flow of a technical conversation. A person may think of a question after a call, then need to send a message and wait, losing the thread of the discussion. In-person pair programming lets two people learn from each other and develop ideas immediately. Joe does not argue for one universal work arrangement. He says the right balance depends on the company, culture, and team, but difficult data problems can benefit from meeting physically and using a whiteboard. Sadie agrees that virtual whiteboards have not matched the value of writing together in the same room.

### Pair programming and empathy are the panel's direct advice
[24:20](https://www.youtube.com/watch?v=A9oLe3bqEpY&t=1460s)
Victor's final advice is to pair program and improve communication between data science and data engineering. He also stresses empathy around time zones and working hours, since colleagues may be several hours ahead or behind each other. Sadie recommends trying the other person's job, or at least learning its terms and constraints. That can make the larger product easier to understand and may even change someone's view of which team they want to join. Joe adds that informal contact matters too. Having lunch with colleagues helps people see one another as people rather than as avatars in Slack.

## Notable quotes
- Demetrios Brinkmann: "I think a lot of times many things can be fixed if we have shared terminology." (01:47)
- Sadie St. Lawrence: "I got into the real world and was like, oh, it's just all SQL and all pipelines." (04:35)
- Joe Reis: "Treating data engineering and data science as separate silos is a good way if you wanted your data projects to ultimately fail." (09:18)
- Sadie St. Lawrence: "I would say whatever side of the fence you're on, do the other person's job, or at least try." (25:25)
- Beverly Wright: "My favorite collaborative tool is the whiteboard." (22:33)

## Tools & references mentioned
- Replit
- Cursor
- OpenAI
- o1
- Claude AI
- ChatGPT
- Fundamentals of Data Engineering
- AWS
- Loom
- Slack

## Who should watch
- You are a data scientist who learned modeling but now has to work with production data, pipelines, or engineering teams.
- Your data engineering and data science groups hand work across a boundary without agreeing on the finished product.
- You manage a distributed technical team and want practical advice on pair programming, shared terminology, and when in-person work helps.

## Related talks

- [Why Data Scientists Should Know Data Engineering](https://mlopstalks.com/talks/why-data-scientists-should-know-data-engineering) (Dan Sullivan, 58:28)
- [Building Better Data Teams](https://mlopstalks.com/talks/building-better-data-teams) (Leanne Fitzpatrick, Financial Times, 1:01:40)
- [All Data Scientists Should Learn Software Engineering Principles](https://mlopstalks.com/talks/all-data-scientists-should-learn-software-engineering-principles) (Catherine Nelson, Freelance Data Scientist, 52:55)
- [A Conversation with Seattle Data Guy](https://mlopstalks.com/talks/a-conversation-with-seattle-data-guy) (Benjamin Rogojan, Seattle Data Guy, 47:11)
- [How to Become a Better Data Scientist: The Definitive Guide](https://mlopstalks.com/talks/how-to-become-a-better-data-scientist-the-definitive-guide) (Alexey Grigorev, OLX Group, 1:00:42)
