T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
T-Mobile Warszawa, Mokotów Mid
Wynagrodzenie do uzgodnienia
🪄 Prompt EngineeringStacjonarnieB2B CONTRACT
Aplikuj na tę ofertę
Wyślemy Twój profil bezpośrednio do firmy.
O roli
## Your responsibilities
- Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.
- Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.
- Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.
- Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.
- Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.
- Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.
- Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.
- Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.
- Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.
- Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.
- Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.
- Maintain audit logging capabilities to support compliance, security investigations, and operational governance.
- Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.
- Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.
- 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.
- At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.
- Strong experience with Kubernetes and OpenShift administration in production environments.
- Proven experience deploying and operating vLLM-based inference platforms in production.
- Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.
- Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
- Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.
- Strong Python programming skills with experience developing automation and operational tooling.
- Experience with Bash scripting and Linux systems administration.
- Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
- Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.
## What we offer
- Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
- You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.
- sharing the costs of sports activities
- private medical care
- sharing the costs of professional training & courses
- life insurance
- remote work opportunities
- flexible working time
- corporate products and services at discounted prices
- mobile phone available for private use
- no dress code
- parking space for employees
- extra social benefits
- sharing the costs of tickets to the movies, theater
- holiday funds
- birthday celebration
- sharing the costs of a streaming platform subscription
- employee referral program
- charity initiatives
- extra leave
- platforma benefitowa
## Recruitment stages
- Prześlij swoje CV
- Spotkaj się z przyszłym liderem/ liderką zespołu
- Witaj w T-Mobile :)
- We are a technology company, and our goal is to create innovative solutions for individual and business clients.
- At T-Mobile, we all live in a magenta world! This color is close to our hearts and means faith in the success of undertaken actions, self-confidence, and endurance.
- That’s who we are as a team.
- At #MagentaTeam , we focus on exchanging experiences, agile work, and quick adaptation to changes! #MagentaTeam is, above all, a mix of different competencies, experiences, personalities, temperaments, and views. And this diversity is our greatest strength.