System Reliability Engineer (Application Support + Automation)

fulcrumdigitalLisbon, PortugalcontractMid Level
Active

Job description

Who are weFulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing. The Role​Plan, manage, and oversee all aspects of a Production Environment Define strategies for Application Performance Monitoring, Optimization in Prod environmentRespond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.Work with a global team spread across tech hubs in multiple geographies and time zones.Ability to share knowledge and explain processes and procedures to others.Able to perform on-call duties on a rotational basis.Occasional off hours work required.RequirementsLinuxShell Scripting ITIL / ITSMSQL - good to haveApplication TroubleshootingAny Monitoring tool (Preferred Splunk/Dynatrace)Jenkins - CI/CDAnsibleGroovy Scripting/YamlGit basic/bit bucketGood To Have​Payments Flows, Switching, Settlements, Authorisation flows.Even Framework architectureWho are weFulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing. The Role​Plan, manage, and oversee all aspects of a Production Environment Define strategies for Application Performance Monitoring, Optimization in Prod environmentRespond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.Work with a global team spread across tech hubs in multiple geographies and time zones.Ability to share knowledge and explain processes and procedures to others.Able to perform on-call duties on a rotational basis.Occasional off hours work required.

Who are weFulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.

Who are weWho are weWho are weWho are weFulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.Fulcrum DigitalFulcrum Digitalis an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.

The Role

The RoleThe RoleThe RoleThe Role​​​​
  • Plan, manage, and oversee all aspects of a Production Environment
Plan, manage, and oversee all aspects of a Production EnvironmentPlan, manage, and oversee all aspects of a Production EnvironmentPlan, manage, and oversee all aspects of a Production Environment
  • Define strategies for Application Performance Monitoring, Optimization in Prod environment
Define strategies for Application Performance Monitoring, Optimization in Prod environmentDefine strategies for Application Performance Monitoring, Optimization in Prod environmentDefine strategies for Application Performance Monitoring, Optimization in Prod environment
  • Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
  • Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.
Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.
  • Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.
Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.
  • Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.
Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.
  • Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
  • Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
  • Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
  • Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.
Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.
  • Work with a global team spread across tech hubs in multiple geographies and time zones.
Work with a global team spread across tech hubs in multiple geographies and time zones.Work with a global team spread across tech hubs in multiple geographies and time zones.Work with a global team spread across tech hubs in multiple geographies and time zones.
  • Ability to share knowledge and explain processes and procedures to others.
Ability to share knowledge and explain processes and procedures to others.Ability to share knowledge and explain processes and procedures to others.Ability to share knowledge and explain processes and procedures to others.
  • Able to perform on-call duties on a rotational basis.
Able to perform on-call duties on a rotational basis.Able to perform on-call duties on a rotational basis.Able to perform on-call duties on a rotational basis.
  • Occasional off hours work required.
Occasional off hours work required.Occasional off hours work required.Occasional off hours work required.