Summary
hackajob is collaborating with Oracle to connect them with exceptional professionals for this role. About the team Join the AI Network Operations team to design, operate, and optimize advanced network systems supporting large-scale AI and cloud infrastructure. As an experienced Network Engineer, you will help operate high-performance RDMA network fabrics powering tier-0 customers in the generative AI industry, while partnering across engineering teams and vendors to ensure performance, reliability, and scalability. Description The OCI AI Infrastructure – Network Operations team operates the high-performance RDMA/RoCE network fabrics powering OCI’s largest AI, GPU, and HPC workloads. These globally distributed networks support some of the most demanding generative AI workloads running on Oracle Cloud Infrastructure. As a Principal Network Engineer, you will lead the design, deployment, operation, and optimization of large-scale RDMA/RoCE network fabrics across OCI’s global cloud infrastructure. You will combine deep networking expertise with strong automation and software engineering skills to improve network performance, scalability, reliability, and operational efficiency. You will design advanced automation, testing, telemetry, and monitoring solutions; lead network validation, incident response, and root cause analysis; and drive performance and capacity improvements across production environments. You will partner with engineering teams, vendors, and customers to resolve complex technical challenges, ensure deployment readiness, and evolve network architecture and operational practices at cloud scale. As a technical leader, you will also mentor engineers, influence architecture and engineering standards, and help shape the tools and systems supporting hundreds of thousands of network devices and millions of servers across OCI. Responsibilities Key Responsibilities Lead the design, deployment, validation, and lifecycle management of large-scale RDMA/RoCE network fabrics supporting OCI AI, GPU, and HPC infrastructure. Translate network architectures into scalable designs and deployment plans, ensuring performance, reliability, and operational readiness across OCI’s global cloud environment. Serve as technical lead for complex network initiatives spanning RDMA/RoCE fabrics, data center networking, automation, testing, deployment, and operations. Develop automation frameworks, tools, scripts, and infrastructure pipelines to improve network deployment, testing, reliability, and operational efficiency. Design test strategies and lead pre-production validation, network change reviews, and deployment readiness for high-performance network fabrics. Build and enhance telemetry, monitoring, dashboards, and alerting to identify network health, congestion, performance, and reliability issues. Lead incident response, complex troubleshooting, root cause analysis, and corrective actions for network issues impacting OCI AI and GPU workloads. Analyze network performance and capacity, including latency, throughput, packet loss, and congestion, to drive scalable improvements. Partner across OCI Network Engineering, SRE, AI Infrastructure, Data Center Operations, product teams, and vendors to deliver reliable network solutions. Mentor engineers and contribute to network architecture, engineering standards, operational tooling, and continuous improvement. Preferred Skills & Experience Experience designing, operating, and troubleshooting large-scale RDMA/RoCE, cloud, data center, or high-performance networks. Knowledge of RDMA, RoCE, Ethernet fabrics, congestion control, QoS, and AI/GPU networking. Expertise in routing and switching technologies, including BGP, OSPF, EVPN-VXLAN, and data center networking. Experience with network automation using Python, Ansible, APIs, or similar technologies. Experience with network telemetry, observability, monitoring, performance analysis, and incident management. Experience supporting hyperscale cloud, AI/GPU, HPC, or large-scale distributed infrastructure. Ability to lead complex technical initiatives and collaborate across engineering, operations, customers, and vendors. Qualifications Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $102,300 to $209,500 per annum. May be eligible for bonus and equity. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9. 11 paid holidays 10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours. 11. Paid parental leave 12. Adoption assistance 13. Employee Stock Purchase Plan 14. Financial planning and group legal 15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted. As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.Job Description
hackajob is collaborating with Oracle to connect them with exceptional professionals for this role. About the team Join the AI Network Operations team to design, operate, and optimize advanced network systems supporting large-scale AI and cloud infrastructure. As an experienced Network Engineer, you will help operate high-performance RDMA network fabrics powering tier-0 customers in the generative AI industry, while partnering across engineering teams and vendors to ensure performance, reliability, and scalability. Description The OCI AI Infrastructure – Network Operations team operates the high-performance RDMA/RoCE network fabrics powering OCI’s largest AI, GPU, and HPC workloads. These globally distributed networks support some of the most demanding generative AI workloads running on Oracle Cloud Infrastructure. As a Principal Network Engineer, you will lead the design, deployment, operation, and optimization of large-scale RDMA/RoCE network fabrics across OCI’s global cloud infrastructure. You will combine deep networking expertise with strong automation and software engineering skills to improve network performance, scalability, reliability, and operational efficiency. You will design advanced automation, testing, telemetry, and monitoring solutions; lead network validation, incident response, and root cause analysis; and drive performance and capacity improvements across production environments. You will partner with engineering teams, vendors, and customers to resolve complex technical challenges, ensure deployment readiness, and evolve network architecture and operational practices at cloud scale. As a technical leader, you will also mentor engineers, influence architecture and engineering standards, and help shape the tools and systems supporting hundreds of thousands of network devices and millions of servers across OCI. Responsibilities Key Responsibilities Lead the design, deployment, validation, and lifecycle management of large-scale RDMA/RoCE network fabrics supporting OCI AI, GPU, and HPC infrastructure. Translate network architectures into scalable designs and deployment plans, ensuring performance, reliability, and operational readiness across OCI’s global cloud environment. Serve as technical lead for complex network initiatives spanning RDMA/RoCE fabrics, data center networking, automation, testing, deployment, and operations. Develop automation frameworks, tools, scripts, and infrastructure pipelines to improve network deployment, testing, reliability, and operational efficiency. Design test strategies and lead pre-production validation, network change reviews, and deployment readiness for high-performance network fabrics. Build and enhance telemetry, monitoring, dashboards, and alerting to identify network health, congestion, performance, and reliability issues. Lead incident response, complex troubleshooting, root cause analysis, and corrective actions for network issues impacting OCI AI and GPU workloads. Analyze network performance and capacity, including latency, throughput, packet loss, and congestion, to drive scalable improvements. Partner across OCI Network Engineering, SRE, AI Infrastructure, Data Center Operations, product teams, and vendors to deliver reliable network solutions. Mentor engineers and contribute to network architecture, engineering standards, operational tooling, and continuous improvement. Preferred Skills & Experience Experience designing, operating, and troubleshooting large-scale RDMA/RoCE, cloud, data center, or high-performance networks. Knowledge of RDMA, RoCE, Ethernet fabrics, congestion control, QoS, and AI/GPU networking. Expertise in routing and switching technologies, including BGP, OSPF, EVPN-VXLAN, and data center networking. Experience with network automation using Python, Ansible, APIs, or similar technologies. Experience with network telemetry, observability, monitoring, performance analysis, and incident management. Experience supporting hyperscale cloud, AI/GPU, HPC, or large-scale distributed infrastructure. Ability to lead complex technical initiatives and collaborate across engineering, operations, customers, and vendors. Qualifications Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $102,300 to $209,500 per annum. May be eligible for bonus and equity. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9. 11 paid holidays 10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours. 11. Paid parental leave 12. Adoption assistance 13. Employee Stock Purchase Plan 14. Financial planning and group legal 15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted. As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.
Government Jobs
Government jobs offer stability, competitive benefits, and the chance to make a meaningful impact on your community and country.
Whether you’re starting your career or seeking new opportunities, these roles provide pathways for growth, security, and service.
Explore positions across a wide range of fields and take the first step toward a rewarding future in public service.
MORE JOBS
-
$
Lead Principal Core Infrastructure Engineer
- Nashville
- Oracle
- Oct 10, 2026
-
$
Principal Core Infrastructure Engineer
- Nashville
- Oracle
- Oct 10, 2026
-
$
Lead Principal Systems Software Engineer
- Nashville
- Oracle
- Oct 10, 2026
-
$
Software Developer 3
- Nashville
- Oracle
- Oct 10, 2026
-
$
Principal Core Infrastructure Engineer
- Nashville
- Oracle
- Oct 10, 2026
-
$
Senior Core Infrastructure Engineer
- Nashville
- Oracle
- Oct 10, 2026