Site reliability engineer
Budget: $750 – $1,500 USD
• Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence.
• Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services.
• Develop designs, architectures, standards, and methods for large-scale distributed systems.
• Design, implement and integrate monitoring solutions to pursue the high reliability.
• Investigate and implement approaches which makes system high alliable and fault tolerant.
• Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning.
• Work with other engineers within the HCGBU – Delivery Platform team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services.
• Articulate technical characteristics of services and technology areas and guide development teams to engineer and add capabilities to internal Oracle services.
• Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).
• Utilize a deep understanding of service topology and the dependencies required to troubleshoot issues and define mitigations.
• Understand and explain the effect of product architecture decisions on distributed systems.
• Serve as part of a 24x7 On Call rotation in support of the HCGBU – Delivery Platform.
• Professional curiosity and a desire to a develop deep understanding of services and technologies
Mandatory Qualifications:
• Experience with Python, bash, and/or other scripting programming
• Experience working with fault tolerant, highly available, high throughput, distributed, scalable systems
• Aptitude to be a good team player and the desire to learn and implement new Cloud technologies as needed
• Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services.
• Develop designs, architectures, standards, and methods for large-scale distributed systems.
• Design, implement and integrate monitoring solutions to pursue the high reliability.
• Investigate and implement approaches which makes system high alliable and fault tolerant.
• Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning.
• Work with other engineers within the HCGBU – Delivery Platform team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services.
• Articulate technical characteristics of services and technology areas and guide development teams to engineer and add capabilities to internal Oracle services.
• Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).
• Utilize a deep understanding of service topology and the dependencies required to troubleshoot issues and define mitigations.
• Understand and explain the effect of product architecture decisions on distributed systems.
• Serve as part of a 24x7 On Call rotation in support of the HCGBU – Delivery Platform.
• Professional curiosity and a desire to a develop deep understanding of services and technologies
Mandatory Qualifications:
• Experience with Python, bash, and/or other scripting programming
• Experience working with fault tolerant, highly available, high throughput, distributed, scalable systems
• Aptitude to be a good team player and the desire to learn and implement new Cloud technologies as needed