Data Scientist
Job description
General Information
Job Title Data Scientist Job ID 108666 Work Areas Analytics, Data & Research, Management Consulting, Product Management & Innovation, Technology & Engineering Employment Type Permanent Full-Time Location(s) Madrid, WarsawJob Title Data Scientist Job ID 108666 Work Areas Analytics, Data & Research, Management Consulting, Product Management & Innovation, Technology & Engineering Employment Type Permanent Full-Time Location(s) Madrid, WarsawJob Title Data ScientistJob TitleData ScientistJob ID 108666Job ID108666Work Areas Analytics, Data & Research, Management Consulting, Product Management & Innovation, Technology & EngineeringWork AreasAnalytics, Data & Research, Management Consulting, Product Management & Innovation, Technology & EngineeringEmployment Type Permanent Full-TimeEmployment TypePermanent Full-TimeLocation(s) Madrid, WarsawLocation(s)Madrid, WarsawDescription & RequirementsDescription & RequirementsDescription & Requirements
WHAT MAKES US A GREAT PLACE TO WORKWe are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times. Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally.WHO YOU’LL WORK WITHAbout Coro℠The Coro℠ business unit brings together Bain’s proprietary suite of software-as-a-service (SaaS) and data-as-a-service (DaaS) tools, including cloud-based software, online capability assessments, and advanced analytics, focused on enabling Commercial Excellence for B2B companies.You’ll work closely with data scientists, data engineers, and software engineers to develop sophisticated approaches to entity resolution at scale. Your focus will be on experimentation and model quality, while your engineering partners will help bring successful approaches into production.WHERE YOU’LL FIT WITHIN THE TEAMAs a Senior Applied Data Scientist, you’ll focus on one of the most challenging problems in large-scale data: determining when records from different sources refer to the same real-world business.You’ll develop and test machine learning, embedding, and large language model (LLM) approaches that improve how we match and resolve complex entity data. You’ll explore how far emerging foundation-model techniques can improve match quality while ensuring solutions remain practical, scalable, and cost-effective across hundreds of millions of entities.This is a highly applied role where you’ll have the opportunity to experiment with emerging AI techniques, measure their impact, and work with engineering teams to turn the strongest ideas into scalable solutions.WHAT YOU’LL DODevelop Smarter Entity MatchingBuild and evaluate machine learning, embedding, and LLM-based approaches for entity resolutionImprove how our systems handle complex and messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchiesDevelop scoring and ranking approaches that distinguish genuine matches from lookalikes, duplicates, and unrelated entitiesEvaluate AI and machine learning techniques while balancing accuracy, scalability, and costDesign solutions for large-scale use, identifying where sophisticated models add value and where more efficient approaches can achieve comparable resultsImprove Experimentation & Model QualityDevelop robust approaches to measuring match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review requirementsHelp build trusted benchmark datasets to compare new approaches with existing matching methods before production rolloutExplore LLM-assisted review and validation both as a matching technique and as a benchmark for more scalable approachesTranslate ambiguous matching challenges into clear hypotheses, experiments, metrics, and recommendationsConduct detailed error analysis to understand model behavior and identify opportunities for improvementBring Successful Approaches to ProductionPartner closely with data and software engineers to turn promising prototypes into production-ready matching solutionsProvide clear model specifications, expected behaviors, evaluation results, edge cases, and rollout criteriaHelp determine the right matching techniques for different data tiers, confidence levels, and cost profilesMeasure impact, diagnose regressions, and recommend improvements to models and matching logicClearly communicate trade-offs across model quality, scale, cost, latency, explainability, and operational riskABOUT YOURequired Qualifications5–8 years of relevant professional experience in applied data science, machine learning, or a related fieldStrong applied machine learning expertise, including hands-on experience building and evaluating models using real-world dataExcellent Python and SQL skillsPractical experience with embeddings, semantic similarity, LLMs, or related AI techniquesHands-on experience training supervised and unsupervised models, including classification and NLP applicationsWorking knowledge of neural networks and transformer architecturesExperience with machine learning frameworks such as TensorFlow, PyTorch, or PyCaretExperience retraining taxonomy classifiers or maintaining classification models in productionStrong experimental judgment, including the ability to define baselines, evaluation metrics, test sets, and error analysesAbility to clearly explain model behavior, technical trade-offs, and edge cases to engineering and business stakeholdersPreferred QualificationsExperience with entity resolution, record linkage, deduplication, or similar matching problemsExperience with ranking, similarity scoring, retrieval, clustering, or candidate-generation techniquesExperience applying LLMs or embeddings to large-scale business problems where performance, scalability, and cost are important considerationsExposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQueryFamiliarity with company, domain, website, firmographic, or other business-entity dataWHAT MAKES US A GREAT PLACE TO WORKWe are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times. Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally.WHO YOU’LL WORK WITHAbout Coro℠The Coro℠ business unit brings together Bain’s proprietary suite of software-as-a-service (SaaS) and data-as-a-service (DaaS) tools, including cloud-based software, online capability assessments, and advanced analytics, focused on enabling Commercial Excellence for B2B companies.You’ll work closely with data scientists, data engineers, and software engineers to develop sophisticated approaches to entity resolution at scale. Your focus will be on experimentation and model quality, while your engineering partners will help bring successful approaches into production.WHERE YOU’LL FIT WITHIN THE TEAMAs a Senior Applied Data Scientist, you’ll focus on one of the most challenging problems in large-scale data: determining when records from different sources refer to the same real-world business.You’ll develop and test machine learning, embedding, and large language model (LLM) approaches that improve how we match and resolve complex entity data. You’ll explore how far emerging foundation-model techniques can improve match quality while ensuring solutions remain practical, scalable, and cost-effective across hundreds of millions of entities.This is a highly applied role where you’ll have the opportunity to experiment with emerging AI techniques, measure their impact, and work with engineering teams to turn the strongest ideas into scalable solutions.WHAT YOU’LL DODevelop Smarter Entity MatchingBuild and evaluate machine learning, embedding, and LLM-based approaches for entity resolutionImprove how our systems handle complex and messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchiesDevelop scoring and ranking approaches that distinguish genuine matches from lookalikes, duplicates, and unrelated entitiesEvaluate AI and machine learning techniques while balancing accuracy, scalability, and costDesign solutions for large-scale use, identifying where sophisticated models add value and where more efficient approaches can achieve comparable resultsImprove Experimentation & Model QualityDevelop robust approaches to measuring match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review requirementsHelp build trusted benchmark datasets to compare new approaches with existing matching methods before production rolloutExplore LLM-assisted review and validation both as a matching technique and as a benchmark for more scalable approachesTranslate ambiguous matching challenges into clear hypotheses, experiments, metrics, and recommendationsConduct detailed error analysis to understand model behavior and identify opportunities for improvementBring Successful Approaches to ProductionPartner closely with data and software engineers to turn promising prototypes into production-ready matching solutionsProvide clear model specifications, expected behaviors, evaluation results, edge cases, and rollout criteriaHelp determine the right matching techniques for different data tiers, confidence levels, and cost profilesMeasure impact, diagnose regressions, and recommend improvements to models and matching logicClearly communicate trade-offs across model quality, scale, cost, latency, explainability, and operational riskABOUT YOURequired Qualifications5–8 years of relevant professional experience in applied data science, machine learning, or a related fieldStrong applied machine learning expertise, including hands-on experience building and evaluating models using real-world dataExcellent Python and SQL skillsPractical experience with embeddings, semantic similarity, LLMs, or related AI techniquesHands-on experience training supervised and unsupervised models, including classification and NLP applicationsWorking knowledge of neural networks and transformer architecturesExperience with machine learning frameworks such as TensorFlow, PyTorch, or PyCaretExperience retraining taxonomy classifiers or maintaining classification models in productionStrong experimental judgment, including the ability to define baselines, evaluation metrics, test sets, and error analysesAbility to clearly explain model behavior, technical trade-offs, and edge cases to engineering and business stakeholdersPreferred QualificationsExperience with entity resolution, record linkage, deduplication, or similar matching problemsExperience with ranking, similarity scoring, retrieval, clustering, or candidate-generation techniquesExperience applying LLMs or embeddings to large-scale business problems where performance, scalability, and cost are important considerationsExposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQueryFamiliarity with company, domain, website, firmographic, or other business-entity dataWHAT MAKES US A GREAT PLACE TO WORKWe are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times. Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally.WHO YOU’LL WORK WITHAbout Coro℠The Coro℠ business unit brings together Bain’s proprietary suite of software-as-a-service (SaaS) and data-as-a-service (DaaS) tools, including cloud-based software, online capability assessments, and advanced analytics, focused on enabling Commercial Excellence for B2B companies.You’ll work closely with data scientists, data engineers, and software engineers to develop sophisticated approaches to entity resolution at scale. Your focus will be on experimentation and model quality, while your engineering partners will help bring successful approaches into production.WHERE YOU’LL FIT WITHIN THE TEAMAs a Senior Applied Data Scientist, you’ll focus on one of the most challenging problems in large-scale data: determining when records from different sources refer to the same real-world business.You’ll develop and test machine learning, embedding, and large language model (LLM) approaches that improve how we match and resolve complex entity data. You’ll explore how far emerging foundation-model techniques can improve match quality while ensuring solutions remain practical, scalable, and cost-effective across hundreds of millions of entities.This is a highly applied role where you’ll have the opportunity to experiment with emerging AI techniques, measure their impact, and work with engineering teams to turn the strongest ideas into scalable solutions.WHAT YOU’LL DODevelop Smarter Entity MatchingBuild and evaluate machine learning, embedding, and LLM-based approaches for entity resolutionImprove how our systems handle complex and messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchiesDevelop scoring and ranking approaches that distinguish genuine matches from lookalikes, duplicates, and unrelated entitiesEvaluate AI and machine learning techniques while balancing accuracy, scalability, and costDesign solutions for large-scale use, identifying where sophisticated models add value and where more efficient approaches can achieve comparable resultsImprove Experimentation & Model QualityDevelop robust approaches to measuring match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review requirementsHelp build trusted benchmark datasets to compare new approaches with existing matching methods before production rolloutExplore LLM-assisted review and validation both as a matching technique and as a benchmark for more scalable approachesTranslate ambiguous matching challenges into clear hypotheses, experiments, metrics, and recommendationsConduct detailed error analysis to understand model behavior and identify opportunities for improvementBring Successful Approaches to ProductionPartner closely with data and software engineers to turn promising prototypes into production-ready matching solutionsProvide clear model specifications, expected behaviors, evaluation results, edge cases, and rollout criteriaHelp determine the right matching techniques for different data tiers, confidence levels, and cost profilesMeasure impact, diagnose regressions, and recommend improvements to models and matching logicClearly communicate trade-offs across model quality, scale, cost, latency, explainability, and operational riskABOUT YOURequired Qualifications5–8 years of relevant professional experience in applied data science, machine learning, or a related fieldStrong applied machine learning expertise, including hands-on experience building and evaluating models using real-world dataExcellent Python and SQL skillsPractical experience with embeddings, semantic similarity, LLMs, or related AI techniquesHands-on experience training supervised and unsupervised models, including classification and NLP applicationsWorking knowledge of neural networks and transformer architecturesExperience with machine learning frameworks such as TensorFlow, PyTorch, or PyCaretExperience retraining taxonomy classifiers or maintaining classification models in productionStrong experimental judgment, including the ability to define baselines, evaluation metrics, test sets, and error analysesAbility to clearly explain model behavior, technical trade-offs, and edge cases to engineering and business stakeholdersPreferred QualificationsExperience with entity resolution, record linkage, deduplication, or similar matching problemsExperience with ranking, similarity scoring, retrieval, clustering, or candidate-generation techniquesExperience applying LLMs or embeddings to large-scale business problems where performance, scalability, and cost are important considerationsExposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQueryFamiliarity with company, domain, website, firmographic, or other business-entity dataWHAT MAKES US A GREAT PLACE TO WORKWe are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times. Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally.WHO YOU’LL WORK WITHAbout Coro℠The Coro℠ business unit brings together Bain’s proprietary suite of software-as-a-service (SaaS) and data-as-a-service (DaaS) tools, including cloud-based software, online capability assessments, and advanced analytics, focused on enabling Commercial Excellence for B2B companies.You’ll work closely with data scientists, data engineers, and software engineers to develop sophisticated approaches to entity resolution at scale. Your focus will be on experimentation and model quality, while your engineering partners will help bring successful approaches into production.WHERE YOU’LL FIT WITHIN THE TEAMAs a Senior Applied Data Scientist, you’ll focus on one of the most challenging problems in large-scale data: determining when records from different sources refer to the same real-world business.You’ll develop and test machine learning, embedding, and large language model (LLM) approaches that improve how we match and resolve complex entity data. You’ll explore how far emerging foundation-model techniques can improve match quality while ensuring solutions remain practical, scalable, and cost-effective across hundreds of millions of entities.This is a highly applied role where you’ll have the opportunity to experiment with emerging AI techniques, measure their impact, and work with engineering teams to turn the strongest ideas into scalable solutions.WHAT YOU’LL DODevelop Smarter Entity MatchingBuild and evaluate machine learning, embedding, and LLM-based approaches for entity resolutionImprove how our systems handle complex and messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchiesDevelop scoring and ranking approaches that distinguish genuine matches from lookalikes, duplicates, and unrelated entitiesEvaluate AI and machine learning techniques while balancing accuracy, scalability, and costDesign solutions for large-scale use, identifying where sophisticated models add value and where more efficient approaches can achieve comparable resultsImprove Experimentation & Model QualityDevelop robust approaches to measuring match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review requirementsHelp build trusted benchmark datasets to compare new approaches with existing matching methods before production rolloutExplore LLM-assisted review and validation both as a matching technique and as a benchmark for more scalable approachesTranslate ambiguous matching challenges into clear hypotheses, experiments, metrics, and recommendationsConduct detailed error analysis to understand model behavior and identify opportunities for improvementBring Successful Approaches to ProductionPartner closely with data and software engineers to turn promising prototypes into production-ready matching solutionsProvide clear model specifications, expected behaviors, evaluation results, edge cases, and rollout criteriaHelp determine the right matching techniques for different data tiers, confidence levels, and cost profilesMeasure impact, diagnose regressions, and recommend improvements to models and matching logicClearly communicate trade-offs across model quality, scale, cost, latency, explainability, and operational riskABOUT YOURequired Qualifications5–8 years of relevant professional experience in applied data science, machine learning, or a related fieldStrong applied machine learning expertise, including hands-on experience building and evaluating models using real-world dataExcellent Python and SQL skillsPractical experience with embeddings, semantic similarity, LLMs, or related AI techniquesHands-on experience training supervised and unsupervised models, including classification and NLP applicationsWorking knowledge of neural networks and transformer architecturesExperience with machine learning frameworks such as TensorFlow, PyTorch, or PyCaretExperience retraining taxonomy classifiers or maintaining classification models in productionStrong experimental judgment, including the ability to define baselines, evaluation metrics, test sets, and error analysesAbility to clearly explain model behavior, technical trade-offs, and edge cases to engineering and business stakeholdersPreferred QualificationsExperience with entity resolution, record linkage, deduplication, or similar matching problemsExperience with ranking, similarity scoring, retrieval, clustering, or candidate-generation techniquesExperience applying LLMs or embeddings to large-scale business problems where performance, scalability, and cost are important considerationsExposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQueryFamiliarity with company, domain, website, firmographic, or other business-entity dataWHAT MAKES US A GREAT PLACE TO WORK
WHAT MAKES US A GREAT PLACE TO WORKWe are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times. Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally.
We are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times. Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally.WHO YOU’LL WORK WITH
WHO YOU’LL WORK WITHAbout Coro℠
About Coro℠The Coro℠ business unit brings together Bain’s proprietary suite of software-as-a-service (SaaS) and data-as-a-service (DaaS) tools, including cloud-based software, online capability assessments, and advanced analytics, focused on enabling Commercial Excellence for B2B companies.
The Coro℠ business unit brings together Bain’s proprietary suite of software-as-a-service (SaaS) and data-as-a-service (DaaS) tools, including cloud-based software, online capability assessments, and advanced analytics, focused on enabling Commercial Excellence for B2B companies.You’ll work closely with data scientists, data engineers, and software engineers to develop sophisticated approaches to entity resolution at scale. Your focus will be on experimentation and model quality, while your engineering partners will help bring successful approaches into production.
You’ll work closely with data scientists, data engineers, and software engineers to develop sophisticated approaches to entity resolution at scale. Your focus will be on experimentation and model quality, while your engineering partners will help bring successful approaches into production.WHERE YOU’LL FIT WITHIN THE TEAM
WHERE YOU’LL FIT WITHIN THE TEAMAs a Senior Applied Data Scientist, you’ll focus on one of the most challenging problems in large-scale data: determining when records from different sources refer to the same real-world business.
As a Senior Applied Data Scientist, you’ll focus on one of the most challenging problems in large-scale data: determining when records from different sources refer to the same real-world business.You’ll develop and test machine learning, embedding, and large language model (LLM) approaches that improve how we match and resolve complex entity data. You’ll explore how far emerging foundation-model techniques can improve match quality while ensuring solutions remain practical, scalable, and cost-effective across hundreds of millions of entities.
You’ll develop and test machine learning, embedding, and large language model (LLM) approaches that improve how we match and resolve complex entity data. You’ll explore how far emerging foundation-model techniques can improve match quality while ensuring solutions remain practical, scalable, and cost-effective across hundreds of millions of entities.This is a highly applied role where you’ll have the opportunity to experiment with emerging AI techniques, measure their impact, and work with engineering teams to turn the strongest ideas into scalable solutions.
This is a highly applied role where you’ll have the opportunity to experiment with emerging AI techniques, measure their impact, and work with engineering teams to turn the strongest ideas into scalable solutions.WHAT YOU’LL DO
WHAT YOU’LL DODevelop Smarter Entity Matching
Develop Smarter Entity Matching- Build and evaluate machine learning, embedding, and LLM-based approaches for entity resolution
- Improve how our systems handle complex and messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies
- Develop scoring and ranking approaches that distinguish genuine matches from lookalikes, duplicates, and unrelated entities
- Evaluate AI and machine learning techniques while balancing accuracy, scalability, and cost
- Design solutions for large-scale use, identifying where sophisticated models add value and where more efficient approaches can achieve comparable results
Improve Experimentation & Model Quality
Improve Experimentation & Model Quality- Develop robust approaches to measuring match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review requirements
- Help build trusted benchmark datasets to compare new approaches with existing matching methods before production rollout
- Explore LLM-assisted review and validation both as a matching technique and as a benchmark for more scalable approaches
- Translate ambiguous matching challenges into clear hypotheses, experiments, metrics, and recommendations
- Conduct detailed error analysis to understand model behavior and identify opportunities for improvement
Bring Successful Approaches to Production
Bring Successful Approaches to Production- Partner closely with data and software engineers to turn promising prototypes into production-ready matching solutions
- Provide clear model specifications, expected behaviors, evaluation results, edge cases, and rollout criteria
- Help determine the right matching techniques for different data tiers, confidence levels, and cost profiles
- Measure impact, diagnose regressions, and recommend improvements to models and matching logic
- Clearly communicate trade-offs across model quality, scale, cost, latency, explainability, and operational risk
ABOUT YOU
ABOUT YOURequired Qualifications
Required Qualifications- 5–8 years of relevant professional experience in applied data science, machine learning, or a related field
- Strong applied machine learning expertise, including hands-on experience building and evaluating models using real-world data
- Excellent Python and SQL skills
- Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques
- Hands-on experience training supervised and unsupervised models, including classification and NLP applications
- Working knowledge of neural networks and transformer architectures
- Experience with machine learning frameworks such as TensorFlow, PyTorch, or PyCaret
- Experience retraining taxonomy classifiers or maintaining classification models in production
- Strong experimental judgment, including the ability to define baselines, evaluation metrics, test sets, and error analyses
- Ability to clearly explain model behavior, technical trade-offs, and edge cases to engineering and business stakeholders
Preferred Qualifications
Preferred Qualifications- Experience with entity resolution, record linkage, deduplication, or similar matching problems
- Experience with ranking, similarity scoring, retrieval, clustering, or candidate-generation techniques
- Experience applying LLMs or embeddings to large-scale business problems where performance, scalability, and cost are important considerations
- Exposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQuery
- Familiarity with company, domain, website, firmographic, or other business-entity data