Software Engineer II (Data Engineering)

zagenoBangalorefull time
Activeverified Jul 25, 2026

Job description

Privacy Policy{"@context":"http://schema.org","@type":"JobPosting","title":"Software Engineer II (Data Engineering)","description":"

About the Role<\/span><\/p>\n

ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.<\/span>The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.<\/span><\/p>\n

You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.<\/span><\/p>\n


<\/p>\n

In this role you will:<\/span><\/p>\n

    \n
  • Own reliability and performance<\/span> of operational pipelines across our product catalog infrastructure.<\/span><\/li>\n
  • Identify and reduce technical debt<\/span>, replacing reactive patches with designed, testable logic.<\/span><\/li>\n
  • Build and maintain low latency data APIs<\/span> that serve downstream operational and analytics consumers.<\/span><\/li>\n
  • Implement CDC patterns<\/span> to keep catalog data synchronized across systems with minimal lag.<\/span><\/li>\n
  • Build monitoring, observability, and automated testing<\/span> so failures surface before stakeholders report them.<\/span><\/li>\n
  • Design and implement unit standardization<\/span> and master data logic at catalog scale.<\/span><\/li>\n
  • Translate business requirements<\/span> from non-technical stakeholders into durable pipeline logic.<\/span><\/li>\n
  • Own code versioning, deployment, and incident response<\/span> for your layer.<\/span><\/li>\n
  • Leverage your expertise<\/span> in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP.<\/span><\/li>\n<\/ul>\n


    <\/p>\n

    About you:<\/span><\/p>\n

    Required:<\/span><\/p>\n

      \n
    • 3+ years as a Data Engineer, including solo or primary ownership of production pipelines<\/span>
      <\/span><\/li>\n
    • Strong Python – data engineering, transformation logic, testing discipline<\/span>
      <\/span><\/li>\n
    • Strong SQL with ability to write correct queries, identify and refactor anti-patterns<\/span>
      <\/span><\/li>\n
    • Databricks, Delta Lake, Airflow for production orchestration<\/span>
      <\/span><\/li>\n
    • Experience with CDC patterns for real-time or near-real-time data synchronization<\/span>
      <\/span><\/li>\n
    • Experience building low latency APIs serving operational or analytical consumers<\/span>
      <\/span><\/li>\n
    • Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought<\/span>
      <\/span><\/li>\n
    • Operates independently under ambiguity; designs systems to be maintained, not just to run<\/span>

      <\/span><\/li>\n<\/ul>\n

      Preferred:<\/span><\/p>\n

        \n
      • Kafka or equivalent event streaming platform experience<\/span>
        <\/span><\/li>\n
      • Experience with entity matching, deduplication, or master data management<\/span>
        <\/span><\/li>\n
      • Exposure to ML pipeline support in production<\/span>
        <\/span><\/li>\n
      • Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition)<\/span>

        <\/span><\/li>\n<\/ul>\n

        What success looks like:<\/span><\/p>\n

          \n
        • Engineers ship data products:<\/span> Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers.<\/span><\/li>\n
        • Reliability:<\/span> Pipelines run reliably with minimal manual intervention.<\/span><\/li>\n
        • Performance:<\/span> Data latency and downtime decrease measurably over time.<\/span><\/li>\n
        • Data Quality:<\/span> Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes.<\/span><\/li>\n
        • Scalability:<\/span> Infrastructure scales with volume growth without proportional cost increase.<\/span><\/li>\n
        • Reduction of Debt:<\/span> Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic.<\/span><\/li>\n
        • Synchronization:<\/span> CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.<\/span><\/li>\n<\/ul>","identifier":{"@type":"PropertyValue","name":"ZAGENO","value":"103"},"datePosted":"2026-07-18","employmentType":"OTHER","hiringOrganization":{"@type":"Organization","name":"ZAGENO","logo":"https://images7.bamboohr.com/423749/logos/cropped.jpg?v=29"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Bangalore","addressRegion":"Karnataka","postalCode":"560037","addressCountry":"India"}},"url":"https://zageno.bamboohr.com/careers/103"}Job OpeningsSoftware Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs. Apply for This JobLink to This JobLocationBangalore, Karnataka (Hybrid)DepartmentMerchandisingEmployment TypeFull-TimeMinimum ExperienceMid-levelPrivacy Policy • Terms of Service • © BambooHR All rights reserved.

          Privacy Policy{"@context":"http://schema.org","@type":"JobPosting","title":"Software Engineer II (Data Engineering)","description":"

          About the Role<\/span><\/p>\n

          ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.<\/span>The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.<\/span><\/p>\n

          You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.<\/span><\/p>\n


          <\/p>\n

          In this role you will:<\/span><\/p>\n

            \n
          • Own reliability and performance<\/span> of operational pipelines across our product catalog infrastructure.<\/span><\/li>\n
          • Identify and reduce technical debt<\/span>, replacing reactive patches with designed, testable logic.<\/span><\/li>\n
          • Build and maintain low latency data APIs<\/span> that serve downstream operational and analytics consumers.<\/span><\/li>\n
          • Implement CDC patterns<\/span> to keep catalog data synchronized across systems with minimal lag.<\/span><\/li>\n
          • Build monitoring, observability, and automated testing<\/span> so failures surface before stakeholders report them.<\/span><\/li>\n
          • Design and implement unit standardization<\/span> and master data logic at catalog scale.<\/span><\/li>\n
          • Translate business requirements<\/span> from non-technical stakeholders into durable pipeline logic.<\/span><\/li>\n
          • Own code versioning, deployment, and incident response<\/span> for your layer.<\/span><\/li>\n
          • Leverage your expertise<\/span> in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP.<\/span><\/li>\n<\/ul>\n


            <\/p>\n

            About you:<\/span><\/p>\n

            Required:<\/span><\/p>\n

              \n
            • 3+ years as a Data Engineer, including solo or primary ownership of production pipelines<\/span>
              <\/span><\/li>\n
            • Strong Python – data engineering, transformation logic, testing discipline<\/span>
              <\/span><\/li>\n
            • Strong SQL with ability to write correct queries, identify and refactor anti-patterns<\/span>
              <\/span><\/li>\n
            • Databricks, Delta Lake, Airflow for production orchestration<\/span>
              <\/span><\/li>\n
            • Experience with CDC patterns for real-time or near-real-time data synchronization<\/span>
              <\/span><\/li>\n
            • Experience building low latency APIs serving operational or analytical consumers<\/span>
              <\/span><\/li>\n
            • Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought<\/span>
              <\/span><\/li>\n
            • Operates independently under ambiguity; designs systems to be maintained, not just to run<\/span>

              <\/span><\/li>\n<\/ul>\n

              Preferred:<\/span><\/p>\n

                \n
              • Kafka or equivalent event streaming platform experience<\/span>
                <\/span><\/li>\n
              • Experience with entity matching, deduplication, or master data management<\/span>
                <\/span><\/li>\n
              • Exposure to ML pipeline support in production<\/span>
                <\/span><\/li>\n
              • Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition)<\/span>

                <\/span><\/li>\n<\/ul>\n

                What success looks like:<\/span><\/p>\n

                  \n
                • Engineers ship data products:<\/span> Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers.<\/span><\/li>\n
                • Reliability:<\/span> Pipelines run reliably with minimal manual intervention.<\/span><\/li>\n
                • Performance:<\/span> Data latency and downtime decrease measurably over time.<\/span><\/li>\n
                • Data Quality:<\/span> Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes.<\/span><\/li>\n
                • Scalability:<\/span> Infrastructure scales with volume growth without proportional cost increase.<\/span><\/li>\n
                • Reduction of Debt:<\/span> Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic.<\/span><\/li>\n
                • Synchronization:<\/span> CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.<\/span><\/li>\n<\/ul>","identifier":{"@type":"PropertyValue","name":"ZAGENO","value":"103"},"datePosted":"2026-07-18","employmentType":"OTHER","hiringOrganization":{"@type":"Organization","name":"ZAGENO","logo":"https://images7.bamboohr.com/423749/logos/cropped.jpg?v=29"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Bangalore","addressRegion":"Karnataka","postalCode":"560037","addressCountry":"India"}},"url":"https://zageno.bamboohr.com/careers/103"}Job OpeningsSoftware Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs. Apply for This JobLink to This JobLocationBangalore, Karnataka (Hybrid)DepartmentMerchandisingEmployment TypeFull-TimeMinimum ExperienceMid-levelPrivacy Policy • Terms of Service • © BambooHR All rights reserved.

                  Privacy Policy{"@context":"http://schema.org","@type":"JobPosting","title":"Software Engineer II (Data Engineering)","description":"

                  About the Role<\/span><\/p>\n

                  ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.<\/span>The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.<\/span><\/p>\n

                  You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.<\/span><\/p>\n


                  <\/p>\n

                  In this role you will:<\/span><\/p>\n

                    \n
                  • Own reliability and performance<\/span> of operational pipelines across our product catalog infrastructure.<\/span><\/li>\n
                  • Identify and reduce technical debt<\/span>, replacing reactive patches with designed, testable logic.<\/span><\/li>\n
                  • Build and maintain low latency data APIs<\/span> that serve downstream operational and analytics consumers.<\/span><\/li>\n
                  • Implement CDC patterns<\/span> to keep catalog data synchronized across systems with minimal lag.<\/span><\/li>\n
                  • Build monitoring, observability, and automated testing<\/span> so failures surface before stakeholders report them.<\/span><\/li>\n
                  • Design and implement unit standardization<\/span> and master data logic at catalog scale.<\/span><\/li>\n
                  • Translate business requirements<\/span> from non-technical stakeholders into durable pipeline logic.<\/span><\/li>\n
                  • Own code versioning, deployment, and incident response<\/span> for your layer.<\/span><\/li>\n
                  • Leverage your expertise<\/span> in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP.<\/span><\/li>\n<\/ul>\n


                    <\/p>\n

                    About you:<\/span><\/p>\n

                    Required:<\/span><\/p>\n

                      \n
                    • 3+ years as a Data Engineer, including solo or primary ownership of production pipelines<\/span>
                      <\/span><\/li>\n
                    • Strong Python – data engineering, transformation logic, testing discipline<\/span>
                      <\/span><\/li>\n
                    • Strong SQL with ability to write correct queries, identify and refactor anti-patterns<\/span>
                      <\/span><\/li>\n
                    • Databricks, Delta Lake, Airflow for production orchestration<\/span>
                      <\/span><\/li>\n
                    • Experience with CDC patterns for real-time or near-real-time data synchronization<\/span>
                      <\/span><\/li>\n
                    • Experience building low latency APIs serving operational or analytical consumers<\/span>
                      <\/span><\/li>\n
                    • Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought<\/span>
                      <\/span><\/li>\n
                    • Operates independently under ambiguity; designs systems to be maintained, not just to run<\/span>

                      <\/span><\/li>\n<\/ul>\n

                      Preferred:<\/span><\/p>\n

                        \n
                      • Kafka or equivalent event streaming platform experience<\/span>
                        <\/span><\/li>\n
                      • Experience with entity matching, deduplication, or master data management<\/span>
                        <\/span><\/li>\n
                      • Exposure to ML pipeline support in production<\/span>
                        <\/span><\/li>\n
                      • Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition)<\/span>

                        <\/span><\/li>\n<\/ul>\n

                        What success looks like:<\/span><\/p>\n

                          \n
                        • Engineers ship data products:<\/span> Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers.<\/span><\/li>\n
                        • Reliability:<\/span> Pipelines run reliably with minimal manual intervention.<\/span><\/li>\n
                        • Performance:<\/span> Data latency and downtime decrease measurably over time.<\/span><\/li>\n
                        • Data Quality:<\/span> Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes.<\/span><\/li>\n
                        • Scalability:<\/span> Infrastructure scales with volume growth without proportional cost increase.<\/span><\/li>\n
                        • Reduction of Debt:<\/span> Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic.<\/span><\/li>\n
                        • Synchronization:<\/span> CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.<\/span><\/li>\n<\/ul>","identifier":{"@type":"PropertyValue","name":"ZAGENO","value":"103"},"datePosted":"2026-07-18","employmentType":"OTHER","hiringOrganization":{"@type":"Organization","name":"ZAGENO","logo":"https://images7.bamboohr.com/423749/logos/cropped.jpg?v=29"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Bangalore","addressRegion":"Karnataka","postalCode":"560037","addressCountry":"India"}},"url":"https://zageno.bamboohr.com/careers/103"}Job OpeningsSoftware Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs. Apply for This JobLink to This JobLocationBangalore, Karnataka (Hybrid)DepartmentMerchandisingEmployment TypeFull-TimeMinimum ExperienceMid-levelPrivacy Policy • Terms of Service • © BambooHR All rights reserved.

                          Privacy Policy{"@context":"http://schema.org","@type":"JobPosting","title":"Software Engineer II (Data Engineering)","description":"

                          About the Role<\/span><\/p>\n

                          ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.<\/span>The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.<\/span><\/p>\n

                          You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.<\/span><\/p>\n


                          <\/p>\n

                          In this role you will:<\/span><\/p>\n

                            \n
                          • Own reliability and performance<\/span> of operational pipelines across our product catalog infrastructure.<\/span><\/li>\n
                          • Identify and reduce technical debt<\/span>, replacing reactive patches with designed, testable logic.<\/span><\/li>\n
                          • Build and maintain low latency data APIs<\/span> that serve downstream operational and analytics consumers.<\/span><\/li>\n
                          • Implement CDC patterns<\/span> to keep catalog data synchronized across systems with minimal lag.<\/span><\/li>\n
                          • Build monitoring, observability, and automated testing<\/span> so failures surface before stakeholders report them.<\/span><\/li>\n
                          • Design and implement unit standardization<\/span> and master data logic at catalog scale.<\/span><\/li>\n
                          • Translate business requirements<\/span> from non-technical stakeholders into durable pipeline logic.<\/span><\/li>\n
                          • Own code versioning, deployment, and incident response<\/span> for your layer.<\/span><\/li>\n
                          • Leverage your expertise<\/span> in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP.<\/span><\/li>\n<\/ul>\n


                            <\/p>\n

                            About you:<\/span><\/p>\n

                            Required:<\/span><\/p>\n

                              \n
                            • 3+ years as a Data Engineer, including solo or primary ownership of production pipelines<\/span>
                              <\/span><\/li>\n
                            • Strong Python – data engineering, transformation logic, testing discipline<\/span>
                              <\/span><\/li>\n
                            • Strong SQL with ability to write correct queries, identify and refactor anti-patterns<\/span>
                              <\/span><\/li>\n
                            • Databricks, Delta Lake, Airflow for production orchestration<\/span>
                              <\/span><\/li>\n
                            • Experience with CDC patterns for real-time or near-real-time data synchronization<\/span>
                              <\/span><\/li>\n
                            • Experience building low latency APIs serving operational or analytical consumers<\/span>
                              <\/span><\/li>\n
                            • Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought<\/span>
                              <\/span><\/li>\n
                            • Operates independently under ambiguity; designs systems to be maintained, not just to run<\/span>

                              <\/span><\/li>\n<\/ul>\n

                              Preferred:<\/span><\/p>\n

                                \n
                              • Kafka or equivalent event streaming platform experience<\/span>
                                <\/span><\/li>\n
                              • Experience with entity matching, deduplication, or master data management<\/span>
                                <\/span><\/li>\n
                              • Exposure to ML pipeline support in production<\/span>
                                <\/span><\/li>\n
                              • Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition)<\/span>

                                <\/span><\/li>\n<\/ul>\n

                                What success looks like:<\/span><\/p>\n

                                  \n
                                • Engineers ship data products:<\/span> Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers.<\/span><\/li>\n
                                • Reliability:<\/span> Pipelines run reliably with minimal manual intervention.<\/span><\/li>\n
                                • Performance:<\/span> Data latency and downtime decrease measurably over time.<\/span><\/li>\n
                                • Data Quality:<\/span> Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes.<\/span><\/li>\n
                                • Scalability:<\/span> Infrastructure scales with volume growth without proportional cost increase.<\/span><\/li>\n
                                • Reduction of Debt:<\/span> Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic.<\/span><\/li>\n
                                • Synchronization:<\/span> CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.<\/span><\/li>\n<\/ul>","identifier":{"@type":"PropertyValue","name":"ZAGENO","value":"103"},"datePosted":"2026-07-18","employmentType":"OTHER","hiringOrganization":{"@type":"Organization","name":"ZAGENO","logo":"https://images7.bamboohr.com/423749/logos/cropped.jpg?v=29"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Bangalore","addressRegion":"Karnataka","postalCode":"560037","addressCountry":"India"}},"url":"https://zageno.bamboohr.com/careers/103"}Job OpeningsSoftware Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs. Apply for This JobLink to This JobLocationBangalore, Karnataka (Hybrid)DepartmentMerchandisingEmployment TypeFull-TimeMinimum ExperienceMid-levelPrivacy Policy • Terms of Service • © BambooHR All rights reserved.

                                  Privacy Policy{"@context":"http://schema.org","@type":"JobPosting","title":"Software Engineer II (Data Engineering)","description":"

                                  About the Role<\/span><\/p>\n

                                  ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.<\/span>The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.<\/span><\/p>\n

                                  You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.<\/span><\/p>\n


                                  <\/p>\n

                                  In this role you will:<\/span><\/p>\n

                                    \n
                                  • Own reliability and performance<\/span> of operational pipelines across our product catalog infrastructure.<\/span><\/li>\n
                                  • Identify and reduce technical debt<\/span>, replacing reactive patches with designed, testable logic.<\/span><\/li>\n
                                  • Build and maintain low latency data APIs<\/span> that serve downstream operational and analytics consumers.<\/span><\/li>\n
                                  • Implement CDC patterns<\/span> to keep catalog data synchronized across systems with minimal lag.<\/span><\/li>\n
                                  • Build monitoring, observability, and automated testing<\/span> so failures surface before stakeholders report them.<\/span><\/li>\n
                                  • Design and implement unit standardization<\/span> and master data logic at catalog scale.<\/span><\/li>\n
                                  • Translate business requirements<\/span> from non-technical stakeholders into durable pipeline logic.<\/span><\/li>\n
                                  • Own code versioning, deployment, and incident response<\/span> for your layer.<\/span><\/li>\n
                                  • Leverage your expertise<\/span> in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP.<\/span><\/li>\n<\/ul>\n


                                    <\/p>\n

                                    About you:<\/span><\/p>\n

                                    Required:<\/span><\/p>\n

                                      \n
                                    • 3+ years as a Data Engineer, including solo or primary ownership of production pipelines<\/span>
                                      <\/span><\/li>\n
                                    • Strong Python – data engineering, transformation logic, testing discipline<\/span>
                                      <\/span><\/li>\n
                                    • Strong SQL with ability to write correct queries, identify and refactor anti-patterns<\/span>
                                      <\/span><\/li>\n
                                    • Databricks, Delta Lake, Airflow for production orchestration<\/span>
                                      <\/span><\/li>\n
                                    • Experience with CDC patterns for real-time or near-real-time data synchronization<\/span>
                                      <\/span><\/li>\n
                                    • Experience building low latency APIs serving operational or analytical consumers<\/span>
                                      <\/span><\/li>\n
                                    • Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought<\/span>
                                      <\/span><\/li>\n
                                    • Operates independently under ambiguity; designs systems to be maintained, not just to run<\/span>

                                      <\/span><\/li>\n<\/ul>\n

                                      Preferred:<\/span><\/p>\n

                                        \n
                                      • Kafka or equivalent event streaming platform experience<\/span>
                                        <\/span><\/li>\n
                                      • Experience with entity matching, deduplication, or master data management<\/span>
                                        <\/span><\/li>\n
                                      • Exposure to ML pipeline support in production<\/span>
                                        <\/span><\/li>\n
                                      • Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition)<\/span>

                                        <\/span><\/li>\n<\/ul>\n

                                        What success looks like:<\/span><\/p>\n

                                          \n
                                        • Engineers ship data products:<\/span> Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers.<\/span><\/li>\n
                                        • Reliability:<\/span> Pipelines run reliably with minimal manual intervention.<\/span><\/li>\n
                                        • Performance:<\/span> Data latency and downtime decrease measurably over time.<\/span><\/li>\n
                                        • Data Quality:<\/span> Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes.<\/span><\/li>\n
                                        • Scalability:<\/span> Infrastructure scales with volume growth without proportional cost increase.<\/span><\/li>\n
                                        • Reduction of Debt:<\/span> Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic.<\/span><\/li>\n
                                        • Synchronization:<\/span> CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.<\/span><\/li>\n<\/ul>","identifier":{"@type":"PropertyValue","name":"ZAGENO","value":"103"},"datePosted":"2026-07-18","employmentType":"OTHER","hiringOrganization":{"@type":"Organization","name":"ZAGENO","logo":"https://images7.bamboohr.com/423749/logos/cropped.jpg?v=29"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Bangalore","addressRegion":"Karnataka","postalCode":"560037","addressCountry":"India"}},"url":"https://zageno.bamboohr.com/careers/103"}Job OpeningsSoftware Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs. Apply for This JobLink to This JobLocationBangalore, Karnataka (Hybrid)DepartmentMerchandisingEmployment TypeFull-TimeMinimum ExperienceMid-level

                                          Privacy Policy

                                          Privacy Policy

                                          Privacy Policy

                                          Privacy Policy

                                          Privacy Policy

                                          Job OpeningsSoftware Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs. Apply for This JobLink to This JobLocationBangalore, Karnataka (Hybrid)DepartmentMerchandisingEmployment TypeFull-TimeMinimum ExperienceMid-level

                                          Job OpeningsSoftware Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.

                                          Job Openings

                                          Job Openings

                                          Job Openings

                                          Software Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)

                                          Software Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)

                                          Software Engineer II (Data Engineering)Merchandising - Bangalore, Karnataka (Hybrid)

                                          Software Engineer II (Data Engineering)

                                          Merchandising - Bangalore, Karnataka (Hybrid)

                                          About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.

                                          About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.

                                          About the Role ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality. You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end. In this role you will: Own reliability and performance of operational pipelines across our product catalog infrastructure. Identify and reduce technical debt, replacing reactive patches with designed, testable logic. Build and maintain low latency data APIs that serve downstream operational and analytics consumers. Implement CDC patterns to keep catalog data synchronized across systems with minimal lag. Build monitoring, observability, and automated testing so failures surface before stakeholders report them. Design and implement unit standardization and master data logic at catalog scale. Translate business requirements from non-technical stakeholders into durable pipeline logic. Own code versioning, deployment, and incident response for your layer. Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP. About you: Required: 3+ years as a Data Engineer, including solo or primary ownership of production pipelines Strong Python – data engineering, transformation logic, testing discipline Strong SQL with ability to write correct queries, identify and refactor anti-patterns Databricks, Delta Lake, Airflow for production orchestration Experience with CDC patterns for real-time or near-real-time data synchronization Experience building low latency APIs serving operational or analytical consumers Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought Operates independently under ambiguity; designs systems to be maintained, not just to run Preferred: Kafka or equivalent event streaming platform experience Experience with entity matching, deduplication, or master data management Exposure to ML pipeline support in production Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition) What success looks like: Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers. Reliability: Pipelines run reliably with minimal manual intervention. Performance: Data latency and downtime decrease measurably over time. Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes. Scalability: Infrastructure scales with volume growth without proportional cost increase. Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic. Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.

                                          About the Role

                                          About the Role

                                          ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.

                                          ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.

                                          The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.

                                          You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.

                                          You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.

                                          In this role you will:

                                          In this role you will:

                                          Own reliability and performance of operational pipelines across our product catalog infrastructure.

                                          Own reliability and performance

                                          of operational pipelines across our product catalog infrastructure.

                                          Identify and reduce technical debt, replacing reactive patches with designed, testable logic.

                                          Identify and reduce technical debt

                                          , replacing reactive patches with designed, testable logic.

                                          Build and maintain low latency data APIs that serve downstream operational and analytics consumers.

                                          Build and maintain low latency data APIs

                                          that serve downstream operational and analytics consumers.

                                          Implement CDC patterns to keep catalog data synchronized across systems with minimal lag.

                                          Implement CDC patterns

                                          to keep catalog data synchronized across systems with minimal lag.

                                          Build monitoring, observability, and automated testing so failures surface before stakeholders report them.

                                          Build monitoring, observability, and automated testing

                                          so failures surface before stakeholders report them.

                                          Design and implement unit standardization and master data logic at catalog scale.

                                          Design and implement unit standardization

                                          and master data logic at catalog scale.