Databricks Architect (Lead Data Platform Architect)

Sr. Client Partner in San FranciscoNew York, United Statesfull timeLead
Active

Job description

Role OverviewWe are looking for a highly skilled Databricks Architect to design, build, and scale enterprise-grade Lakehouse data platforms. This role will drive architecture strategy, platform standardization, and enterprise data modernization initiatives, leveraging Databricks and cloud ecosystems.The ideal candidate brings deep expertise in Spark, Delta Lake, and cloud-native architecture, along with strong leadership in driving large-scale data transformations.Key ResponsibilitiesData Platform ArchitectureDefine and implement end-to-end Databricks Lakehouse architecture.Design scalable systems for: Batch & real-time data processingStructured & unstructured workloadsEstablish medallion architecture (Bronze, Silver, Gold layers) as a standard.Databricks Platform LeadershipLead deployment and optimization of: Azure Databricks / AWS Databricks / GCP DatabricksDefine standards for: Workspace design & cluster strategyJob orchestrationData storage (Delta Lake)Drive adoption of: Unity CatalogMLflowDatabricks SQL & PhotonSolution Design & EngineeringArchitect robust data ingestion frameworks: Batch (ADF, Airflow)Streaming (Kafka, Event Hub)Define reusable patterns for: ETL/ELT pipelinesData modeling (star schema, data vault, dimensional models)Guide engineering teams on best practices in Spark/PySpark optimization.Performance & Cost OptimizationOptimize workloads for: Query performanceCluster utilizationStorage efficiencyImplement cost governance strategies (auto-scaling, job clusters, spot instances).Data Governance & SecurityArchitect enterprise-grade governance frameworks: Data lineage, cataloging, metadata managementFine-grained access control (RBAC/ABAC)Ensure compliance with data privacy and regulatory standards.Cloud & Ecosystem IntegrationIntegrate Databricks with: Data sources (ERP, CRM, APIs, IoT)BI tools (Power BI, Tableau)ML pipelines and AI platformsCollaborate with cloud architects for: Networking, security, and storage strategies.Leadership & MentorshipProvide architectural guidance to data engineers, scientists, and TPMs.Conduct design reviews and enforce architecture governance.Mentor teams on emerging patterns: Data MeshDataOps / MLOpsGenAI workloads on DatabricksSkills & QualificationsMandatory Skills12+ years of experience in data engineering, architecture, or platform design.5+ years of hands-on experience with: Databricks (must-have)Apache Spark / PySpark / SQLStrong expertise in: Delta LakeDistributed data processingExperience with at least one cloud: Azure (preferred), AWS, or GCPRole OverviewWe are looking for a highly skilled Databricks Architect to design, build, and scale enterprise-grade Lakehouse data platforms. This role will drive architecture strategy, platform standardization, and enterprise data modernization initiatives, leveraging Databricks and cloud ecosystems.The ideal candidate brings deep expertise in Spark, Delta Lake, and cloud-native architecture, along with strong leadership in driving large-scale data transformations.Key ResponsibilitiesData Platform ArchitectureDefine and implement end-to-end Databricks Lakehouse architecture.Design scalable systems for: Batch & real-time data processingStructured & unstructured workloadsEstablish medallion architecture (Bronze, Silver, Gold layers) as a standard.Databricks Platform LeadershipLead deployment and optimization of: Azure Databricks / AWS Databricks / GCP DatabricksDefine standards for: Workspace design & cluster strategyJob orchestrationData storage (Delta Lake)Drive adoption of: Unity CatalogMLflowDatabricks SQL & PhotonSolution Design & EngineeringArchitect robust data ingestion frameworks: Batch (ADF, Airflow)Streaming (Kafka, Event Hub)Define reusable patterns for: ETL/ELT pipelinesData modeling (star schema, data vault, dimensional models)Guide engineering teams on best practices in Spark/PySpark optimization.Performance & Cost OptimizationOptimize workloads for: Query performanceCluster utilizationStorage efficiencyImplement cost governance strategies (auto-scaling, job clusters, spot instances).Data Governance & SecurityArchitect enterprise-grade governance frameworks: Data lineage, cataloging, metadata managementFine-grained access control (RBAC/ABAC)Ensure compliance with data privacy and regulatory standards.Cloud & Ecosystem IntegrationIntegrate Databricks with: Data sources (ERP, CRM, APIs, IoT)BI tools (Power BI, Tableau)ML pipelines and AI platformsCollaborate with cloud architects for: Networking, security, and storage strategies.Leadership & MentorshipProvide architectural guidance to data engineers, scientists, and TPMs.Conduct design reviews and enforce architecture governance.Mentor teams on emerging patterns: Data MeshDataOps / MLOpsGenAI workloads on DatabricksSkills & QualificationsMandatory Skills12+ years of experience in data engineering, architecture, or platform design.5+ years of hands-on experience with: Databricks (must-have)Apache Spark / PySpark / SQLStrong expertise in: Delta LakeDistributed data processingExperience with at least one cloud: Azure (preferred), AWS, or GCPRole OverviewWe are looking for a highly skilled Databricks Architect to design, build, and scale enterprise-grade Lakehouse data platforms. This role will drive architecture strategy, platform standardization, and enterprise data modernization initiatives, leveraging Databricks and cloud ecosystems.The ideal candidate brings deep expertise in Spark, Delta Lake, and cloud-native architecture, along with strong leadership in driving large-scale data transformations.Key ResponsibilitiesData Platform ArchitectureDefine and implement end-to-end Databricks Lakehouse architecture.Design scalable systems for: Batch & real-time data processingStructured & unstructured workloadsEstablish medallion architecture (Bronze, Silver, Gold layers) as a standard.Databricks Platform LeadershipLead deployment and optimization of: Azure Databricks / AWS Databricks / GCP DatabricksDefine standards for: Workspace design & cluster strategyJob orchestrationData storage (Delta Lake)Drive adoption of: Unity CatalogMLflowDatabricks SQL & PhotonSolution Design & EngineeringArchitect robust data ingestion frameworks: Batch (ADF, Airflow)Streaming (Kafka, Event Hub)Define reusable patterns for: ETL/ELT pipelinesData modeling (star schema, data vault, dimensional models)Guide engineering teams on best practices in Spark/PySpark optimization.Performance & Cost OptimizationOptimize workloads for: Query performanceCluster utilizationStorage efficiencyImplement cost governance strategies (auto-scaling, job clusters, spot instances).Data Governance & SecurityArchitect enterprise-grade governance frameworks: Data lineage, cataloging, metadata managementFine-grained access control (RBAC/ABAC)Ensure compliance with data privacy and regulatory standards.Cloud & Ecosystem IntegrationIntegrate Databricks with: Data sources (ERP, CRM, APIs, IoT)BI tools (Power BI, Tableau)ML pipelines and AI platformsCollaborate with cloud architects for: Networking, security, and storage strategies.Leadership & MentorshipProvide architectural guidance to data engineers, scientists, and TPMs.Conduct design reviews and enforce architecture governance.Mentor teams on emerging patterns: Data MeshDataOps / MLOpsGenAI workloads on DatabricksSkills & QualificationsMandatory Skills12+ years of experience in data engineering, architecture, or platform design.5+ years of hands-on experience with: Databricks (must-have)Apache Spark / PySpark / SQLStrong expertise in: Delta LakeDistributed data processingExperience with at least one cloud: Azure (preferred), AWS, or GCP

Role Overview

Role Overview

We are looking for a highly skilled Databricks Architect to design, build, and scale enterprise-grade Lakehouse data platforms. This role will drive architecture strategy, platform standardization, and enterprise data modernization initiatives, leveraging Databricks and cloud ecosystems.

highly skilled Databricks ArchitectLakehouse data platformsarchitecture strategy, platform standardization, and enterprise data modernization initiatives

The ideal candidate brings deep expertise in Spark, Delta Lake, and cloud-native architecture, along with strong leadership in driving large-scale data transformations.

deep expertise in Spark, Delta Lake, and cloud-native architecture

Key Responsibilities

Key Responsibilities

Data Platform Architecture

Data Platform Architecture
  • Define and implement end-to-end Databricks Lakehouse architecture.
end-to-end Databricks Lakehouse architecture
  • Design scalable systems for: Batch & real-time data processingStructured & unstructured workloads
  • Batch & real-time data processing
  • Structured & unstructured workloads
  • Establish medallion architecture (Bronze, Silver, Gold layers) as a standard.
medallion architecture

Databricks Platform Leadership

Databricks Platform Leadership
  • Lead deployment and optimization of: Azure Databricks / AWS Databricks / GCP Databricks
  • Azure Databricks / AWS Databricks / GCP Databricks
  • Define standards for: Workspace design & cluster strategyJob orchestrationData storage (Delta Lake)
  • Workspace design & cluster strategy
  • Job orchestration
  • Data storage (Delta Lake)
  • Drive adoption of: Unity CatalogMLflowDatabricks SQL & Photon
  • Unity Catalog
  • MLflow
  • Databricks SQL & Photon

Solution Design & Engineering

Solution Design & Engineering
  • Architect robust data ingestion frameworks: Batch (ADF, Airflow)Streaming (Kafka, Event Hub)
data ingestion frameworks
  • Batch (ADF, Airflow)
  • Streaming (Kafka, Event Hub)
  • Define reusable patterns for: ETL/ELT pipelinesData modeling (star schema, data vault, dimensional models)
  • ETL/ELT pipelines
  • Data modeling (star schema, data vault, dimensional models)
  • Guide engineering teams on best practices in Spark/PySpark optimization.
best practices in Spark/PySpark optimization

Performance & Cost Optimization

Performance & Cost Optimization
  • Optimize workloads for: Query performanceCluster utilizationStorage efficiency
  • Query performance
  • Cluster utilization
  • Storage efficiency
  • Implement cost governance strategies (auto-scaling, job clusters, spot instances).
cost governance strategies

Data Governance & Security

Data Governance & Security
  • Architect enterprise-grade governance frameworks: Data lineage, cataloging, metadata managementFine-grained access control (RBAC/ABAC)
  • Data lineage, cataloging, metadata management
  • Fine-grained access control (RBAC/ABAC)
  • Ensure compliance with data privacy and regulatory standards.
data privacy and regulatory standards

Cloud & Ecosystem Integration

Cloud & Ecosystem Integration
  • Integrate Databricks with: Data sources (ERP, CRM, APIs, IoT)BI tools (Power BI, Tableau)ML pipelines and AI platforms
  • Data sources (ERP, CRM, APIs, IoT)
  • BI tools (Power BI, Tableau)
  • ML pipelines and AI platforms
  • Collaborate with cloud architects for: Networking, security, and storage strategies.
  • Networking, security, and storage strategies.

Leadership & Mentorship

Leadership & Mentorship
  • Provide architectural guidance to data engineers, scientists, and TPMs.
data engineers, scientists, and TPMs
  • Conduct design reviews and enforce architecture governance.
architecture governance
  • Mentor teams on emerging patterns: Data MeshDataOps / MLOpsGenAI workloads on Databricks
  • Data Mesh
  • DataOps / MLOps
  • GenAI workloads on Databricks

Skills & Qualifications

Skills & Qualifications

Mandatory Skills

Mandatory Skills
  • 12+ years of experience in data engineering, architecture, or platform design.
data engineering, architecture, or platform design
  • 5+ years of hands-on experience with: Databricks (must-have)Apache Spark / PySpark / SQL
  • Databricks (must-have)
Databricks (must-have)
  • Apache Spark / PySpark / SQL
  • Strong expertise in: Delta LakeDistributed data processing
  • Delta Lake
  • Distributed data processing
  • Experience with at least one cloud: Azure (preferred), AWS, or GCP
  • Azure (preferred), AWS, or GCP
Azure (preferred), AWS, or GCPI'm interested{{criteriaIndex}} {{field}} {{captialize(condition)}} ( ) {{cxPropHelpText}} {{cxPropField.consent_details.description}} {{record.Posting_Title}} {{record.Posting_Title}} {{record.Posting_Title}} {{candidate.Email}} {{topMessage}} Job Details {{topMessage}} previous next {{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}} {{message}} Step {{curStepInMandatorySecPrompt}}/{{totalNumOfStepsInMandatorySecPrompt}} For privacy and security purposes, please go through the following points and provide consent. Accept Decline {{getI18n("zr.careers.publicpage.meta.joblisting")}} {{getI18n("zr.careers.publicpage.meta.joblisting")}} {{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}} {{getI18n("zr.cw.autofill.head")}} {{cxPropBodyMessage}} {{getI18n("zr.careers.autopopulate.success")}} {{getI18n("zr.cw.apply.sign",cxPropCompanyInfo.name)}} {{if(cxPropDetails.applyWithSeekButton,cxPropDetails.applyWithSeekButton.buttonLabel,getI18n("zr.quickapply.apply.toggle","Seek"))}} {{getI18n('zr.eeo.questionnaire.portal.maintitle')}} {{botName}} {{getI18n("zr.zia.sb.consent.msg1")}}{{criteriaIndex}}{{criteriaIndex}}{{field}} {{captialize(condition)}}{{field}}{{captialize(condition)}}( )(){{cxPropHelpText}}{{cxPropHelpText}}{{cxPropField.consent_details.description}}

{{cxPropField.consent_details.description}}

{{record.Posting_Title}}{{record.Posting_Title}}{{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}{{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}
  • {{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}{{candidate.Email}}{{candidate.Email}}{{candidate.Email}}

{{candidate.Email}}

{{topMessage}}Job DetailsJob DetailsJob DetailsJob Details{{topMessage}}previousnext{{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}}{{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}}{{title}} {{score}}%

{{title}}

{{title}}{{score}}%Job Description {{unescape(sanitizeHTML(descriptionHTML))}}

Job Description

{{message}}

Step {{curStepInMandatorySecPrompt}}/{{totalNumOfStepsInMandatorySecPrompt}}

For privacy and security purposes, please go through the following points and provide consent.Accept DeclineAccept DeclineAcceptDecline{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}
  • {{getI18n("zr.careers.publicpage.meta.joblisting")}}
{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}
  • {{getI18n("zr.careers.publicpage.meta.joblisting")}}
{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}}{{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}}{{cxPropAssessment.name}}

{{getI18n("zr.questionnaire.question.information.heading")}}

{{getI18n("zr.cw.autofill.head")}} {{cxPropBodyMessage}} {{getI18n("zr.careers.autopopulate.success")}}

{{getI18n("zr.cw.autofill.head")}}

{{cxPropBodyMessage}}

{{getI18n("zr.careers.autopopulate.success")}}

{{getI18n("zr.cw.apply.sign",cxPropCompanyInfo.name)}}

{{getI18n('zr.eeo.questionnaire.portal.maintitle')}}{{botName}}{{botName}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}

Similar jobs