
KYLG AB · Sollentuna
AI Laboratory Infrastructure & Operations Engineer About Us We are an AI and Data Mining company developing. Our AI infrastructure includes NVIDIA GPU clusters,...
AI Laboratory Infrastructure & Operations Engineer
About Us
We are an AI and Data Mining company developing. Our AI infrastructure includes NVIDIA GPU clusters, enterprise AI platforms, large-scale model training and inference environments, and intelligent AI workflow automation.
We are expanding our AI Laboratory and are looking for an experienced engineer to build and operate our AI computing environment.
Responsibilities
You will participate in the planning, deployment, operation, and maintenance of our AI laboratory infrastructure, including:
Design, deploy, and maintain AI server and GPU infrastructure
Install, rack, cable, commission, and maintain AI servers in data center environments
Plan and execute AI infrastructure expansion, upgrades, and hardware lifecycle management
Deploy and manage Linux servers
Install, configure, and maintain NVIDIA CUDA, Docker, Kubernetes, and the AI software stack
Deploy and optimize AI model training and inference environments
Build and maintain AI agent platforms
Develop AI workflows and automation pipelines
Maintain AI computing clusters and GPU infrastructure
Monitor system performance and optimize GPU utilization
Monitor power distribution, cooling systems, and overall hardware health
Troubleshoot AI server hardware, GPU, networking, storage, operating system, and software issues
Support AI researchers and software engineers
Requirements
Required
Strong Linux system administration experience
Experience with AI server hardware deployment, installation, maintenance, and troubleshooting
Experience with data center infrastructure, including power distribution, rack installation, cabling, cooling, and hardware commissioning
Experience deploying and maintaining NVIDIA GPU servers
Experience with Docker
Experience with Python
Experience with CUDA
Experience with AI frameworks (PyTorch / TensorFlow)
Familiar with LLM deployment
Experience with Git
Experience with enterprise networking (10/25/100Gb Ethernet or InfiniBand)
Experience assembling, upgrading and repairing enterprise servers
Preferred
Kubernetes
Slurm
Ansible
GPU cluster management
AI Agent Platforms
LangChain
MCP
Vector Database
RAG
Workflow Automation
Nice to Have
Multi-GPU training
Distributed AI computing
AI Data Center
Enterprise AI Infrastructure
InfiniBand or other high-speed networking
Storage systems (NAS / SAN)
OpenAI API
DeepSeek
Qwen
Llama
We Offer
Competitive salary
Flexible working environment
Opportunity to work with cutting-edge AI technologies
Career development in AI infrastructure and intelligent systems
Language
English required
Chinese is a strong advantage(Chinese (Mandarin) proficiency is highly preferred, as the role involves collaboration with our Chinese engineering teams and technical documentation.)
Xensam: Join the Future of SAM Xensam is the leader in AI-powered, cloud-based Software Asset Management. Our technology brings clarity to complex IT environments, helping users make smart, data-driven decisions and maximize software ROI. Recognized with the Highest Growth Award and ranked #3 Overall Champion at the Main Software 50 Awards Nordics, we’re scaling fast and looking for people who want to join the journey. At Xensam, you play a key role in a team built on energy, focus, and positivity. We value experience, but even more, the person behind it. Together, we build the future of SAM. About the role We're looking for an AI Enablement Engineer to help make AI a natural part of how we work at Xensam. In this role, you'll partner with teams across the business to identify opportunities where AI and automation can create real value. You'll design, build, and maintain AI-powered solutions that simplify workflows, improve productivity, and help our teams focus on what matters most. You'll work across Marketing, Sales, HR, Finance, Operations, Product, and Engineering, combining technical expertise with business understanding to turn ideas into practical solutions. This is a hands-on role with the opportunity to shape how AI is adopted across the entire company. Responsibilities Build and maintain AI-powered solutions and automations tailored to each department's needs, from document processing and customer support to internal knowledge management Drive AI adoption across the company by being accessible, patient, and practical. Help teams move from unfamiliar with AI to confident users Design and develop reusable AI components, prompts, and workflows that different departments can leverage Experiment, iterate, and evaluate new models, frameworks, and approaches. Stay ahead of the curve so we're not playing catch-up Automate intelligently by knowing when to use deterministic, repeatable processes and when to lean on AI for non-deterministic workflows and analytics Build and manage self-hosted AI infrastructure including open-source models (Llama, Qwen, etc.), inference APIs, RAG pipelines, vector databases, and the rest of the stack Build AI solutions on top of existing platforms like ChatGPT, Codex, and other commercial APIs, integrating them into internal workflows and tools Educate and mentor colleagues, running workshops/hackathons and creating documentation that makes AI accessible to everyone, regardless of technical background Partner with every department (operations, HR, finance, marketing, etc.) to understand their workflows and identify where AI or automations can add value Qualifications 3+ years of software development experience Strong experience with Python and API integrations Experience building solutions with LLMs and modern AI frameworks Experience working with RAG, vector databases, or agent frameworks Experience integrating commercial AI platforms such as ChatGPT, Codex, or similar Experience working with open-source AI models and inference platforms Comfortable working with Linux, Docker, and self-hosted environments Strong problem-solving skills and a passion for building practical solutions Excellent communication skills and the ability to collaborate with both technical and non-technical teams Fluent in English, both written and spoken Bonus points if you have Experience with workflow automation platforms (e.g. n8n, Make, Temporal) Experience with cloud AI platforms (AWS Bedrock, Azure AI, GCP Vertex AI) Experience with MLOps or deploying machine learning systems into production Experience fine-tuning or hosting open-source language models Contributions to open-source projects or a portfolio of AI side projects What you get A generous work culture with free drinks and snacks, office massages, and more. Three days in-office (with remote work on Mondays and Fridays). An opportunity to shape your career growth while contributing to the company’s success. A dynamic position embracing "freedom under responsibility". If sales targets are met, all employees enjoy an annual destination trip. Other location-specific benefits. Our values at Xensam Rebellious We challenge the norm and act with initiative – always with responsibility. Humane We foster a caring, inclusive environment that values diversity and respects individuality. Harmony We value balance and create a supportive workplace where people thrive. As part of our recruitment process, we conduct background checks on final candidates to fulfill our commitments to customers and ensure a safe work environment.
Education is one of the most important systems in the world. At Kognity, we build the technology that makes it work better, for the teachers delivering it and the students depending on it. The problems are real, the work is challenging, and the results show up in classrooms every day. We're a 125-person EdTech scale-up powering learning in 120+ countries, helping students and teachers thrive through an intelligent platform that combines rich, interactive pedagogy with smart AI and data. Why join Kognity? Work on problems that matter - Your work directly influences the lives of teachers and students in over 100 countries. The scale is global, and the outcomes are tangible. High ownership, high expectations - You are trusted to take initiative, make decisions and drive outcomes. Responsibility comes early, accountability is real, and results matter. A fast-moving, high-performing team - You will work with smart, driven colleagues across the globe on complex problems. Standards and expectations are high, feedback is direct, and the pace is fast. Continuous growth is the baseline - Everyone is expected and supported to learn quickly, improve constantly and raise their own bar. If you enjoy responsibility, momentum and meaningful challenge, you will thrive here. What you'll do: Take full ownership of the engineering operating system within your first three months: tools, vendor contracts, on-call staffing, and platform support, so no escalations land on the CTO by default and engineers never lose time to broken tooling or unclear process. Audit and overhaul the hiring process end to end, improving interviewer calibration, role consistency, and onboarding so that time-to-hire and time-to-first-meaningful-contribution both come down measurably. Own TLT facilitation and drive EM layer consistency, with sessions that end in decisions, shared practices aligned across teams, and cross-EM escalations resolving below CTO level within six months. Build and maintain the engineering progression framework and performance data infrastructure, from a sharp draft in the first three months to a live, Tech 2.0-aligned framework with salary bands agreed with HR and EMs actively using it in performance conversations. Drive cultural and strategic change from decision to operating reality, particularly the Tech 2.0 rollout, owning the rollout plan, adoption tracking, and the rituals that make change stick across the org over an 18-month horizon. What we're looking for: A background as a senior, principal, or staff engineer, with subsequent experience in engineering leadership or management, ideally including people management in a scaling product organisation. Proven ability to operate without formal authority, driving change through credibility and relationships rather than positional power, and comfortable influencing senior engineers and EMs alike. Strong systems thinking with a bias for simplification: able to look at a complex set of processes and find the version that removes friction without losing what matters. High signal-to-noise communication, both written and verbal, with the ability to translate between strategic direction and operational reality and surface the right information at the right level. A passion for AI and a drive to experiment with new tools to enhance creativity, decisions, and execution. How we hire Discovery call with a Recruiter Hiring manager discussion Case study Values discussion Leadership talk References Every qualified person will be evaluated regardless of age, gender, identity, nationality, ethnicity, sexual orientation, disability status or religion. We're committed to building a diverse, inclusive team and welcome people of all backgrounds, experiences, perspectives, and abilities.
About Keyfactor Our mission is to securely connect the world: humans, machines, and AI. Keyfactor is the leader in trust infrastructure for AI and machines, helping the world’s largest enterprises and government agencies take control of the cryptographic identities that safeguard every digital interaction. Behind the platform is a global team of people who care deeply about the work and each other. We move fast, think big, and show up for one another every day. If you’re looking for work that matters and a team that brings out your best, we hope you’ll trust your future with Keyfactor! Title: Senior Cloud Engineer Location: Stockholm, Sweden Experience: Senior Job Function: Cloud Operations Employment Type: Full-Time Industry: Computer and Network Security About the position The Senior Cloud Engineer designs, implements, and evolves complex cloud infrastructure and automation solutions with a high degree of autonomy. This role owns systems and solutions that have significant impact on service reliability, security posture, and operational scalability. Senior Cloud Engineers act as technical leaders within the team, contributing to standards, mentoring peers, and supporting cross-functional initiatives. The position can based in our Stockholm, we follow a hybrid work model with excellent flexible working practices. Applicants must hold a valid Right to Work in Sweden. Responsibilities Infrastructure as Code (IaC) Owns and evolves complex Infrastructure as Code solutions across the cloud environment Makes independent design decisions within established architectural standards Delivers improvements that enhance scalability, reliability, and team efficiency Provides technical guidance and review for less senior engineers Configuration Designs and maintains complex configuration management solutions Defines configuration patterns and best practices used across the environment Reviews and guides configuration work performed by other engineers Data Center Operations Maintains an accurate and up-to-date inventory of all physical hardware within the data center, such as HSMs, USBs, Smartcards, SSDs and Shock-Proof cases. Informs supervisor when certain materials are running low. Maintains detailed records of all data center operation activities, including changes to tamper evident bag numbers, incidents, and access to client materials. Follows proper procedures on strict guidelines for Root Signings, Disaster Recovery, CRLs and CSR signings. Ensures proper asset disposal procedures are followed. Is able to follow documented procedures on how to properly handle sophisticated hardware security modules and KVM devices. Can perform configurations of the hardware security module and KVM with supervision from more senior peers. Task Automation Designs and builds advanced automation solutions using scripting, APIs, and Azure-native services Drives automation initiatives that improve operational consistency and efficiency Supports process improvement through tooling and automation Networking Designs and implements cloud networking solutions within approved security and compliance constraints Applies deep knowledge of networking protocols and standards to support hosted services Participates in compliance audits and provides technical input Monitoring & Alerting Defines and improves monitoring, alerting, and observability strategies Ensures monitoring logic supports service reliability and rapid incident detection Mentors peers on monitoring best practices Patch Management Leads patch management strategies and execution Resolves escalated patching issues and drives process improvements Contributes to the evolution of patching standards and practices PKI Knowledge Applies strong understanding of PKI concepts in support of customer and hosted environments Performs PKI tasks for custom deployments Contributes to strategy and design discussions for hosted PKI services Produces technical documentation for PKI designs and processes Participates in cross-functional certificate reviews Product Knowledge Serves as a subject matter expert on Keyfactor’s hosted products Troubleshoots complex product issues across hosted environments Participates in strategy and design discussions for new product offerings Supports testing of new product versions and provides feedback to product teams Interfaces with customers on hosted architecture and security posture when required Minimum Qualifications, Education, and Skills High School Diploma or equivalent experience 5+ years of related work experience and strong knowledge within their functional area Expert-level knowledge of Microsoft Azure, including native services and tooling Deep hands-on experience with Infrastructure as Code (IaC) and configuration management tools Advanced proficiency in PowerShell scripting (high complexity) Extensive experience with Windows systems administration, including Windows Roles & Features and Active Directory Advanced problem-solving skills, with the ability to troubleshoot and resolve complex issues Strong written and verbal communication skills with a professional demeanor Strong organizational, time-management, and attention-to-detail skills Proven ability to collaborate effectively within and across cross-functional teams Self-motivated, able to manage complex work and projects to completion with minimal oversight Comfortable working in a fast-paced, deadline-driven environment Travel Requirements Up to 15% travel time required Compensation Salary will be commensurate with experience.