Seoul, South Korea
1 month ago

Job Overview

Pay
Not disclosed

Job description

Location:
Seoul, South Korea
Work arrangement:
On-site

Role Summary

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon.

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.

Designs and implements the low-level runtime stack that drives FuriosaAI's NPU hardware to its theoretical limits — from device driver interfaces and DMA-based I/O to kernel execution scheduling, multi-node inference, and embedded firmware.

Responsibilities

  • Develops the low-level runtime responsible for DMA-based I/O operations and kernel execution scheduling, maximizing inference throughput while minimizing end-to-end latency.
  • Builds and optimizes asynchronous execution pipelines that orchestrate data movement and compute across the NPU hardware.
  • Enables multi-node inference by implementing foundational communication primitives, including RDMA-based data transfer for low-latency, high-bandwidth inter-node operations.
  • Develops embedded firmware (PERT) that runs on the NPU's integrated ARM core, managing on-device scheduling, synchronization, and hardware resource control.
  • Profiles and tunes system-level performance across the full runtime stack — from firmware to user-space — to eliminate bottlenecks in real-world inference workloads.

At Furiosa, you will

  • Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.
  • Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.
  • Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.
  • Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.
  • Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition.

Requirements

  • BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 5+ years of relevant industry experience or equivalent practical experience in systems programming using Rust, C, or C++
  • Solid understanding of computer architecture fundamentals, including memory hierarchy, cache coherency, operating systems, DMA, interrupts, and MMIO
  • Strong communication skills, with the ability to gather requirements and drive technical alignment across teams

Preferred Qualifications

  • Deep expertise in low-latency runtime systems, embedded firmware development, or high-performance I/O — especially in the context of accelerator hardware.
  • Experience designing and implementing low-latency asynchronous execution models and scheduling systems.
  • Experience with DMA engines, scatter-gather I/O, or other zero-copy data transfer mechanisms.
  • Experience developing embedded firmware for ARM-based processors (bare-metal or lightweight RTOS environments).
  • Familiarity with RDMA technologies and high-performance networking for distributed or multi-node systems.
  • Experience with CUDA low-level runtime internals such as CUDA Graphs, stream-based execution, and asynchronous kernel launch optimization.
  • Experience with kernel-level performance optimizations (e.g., Linux kernel modules, eBPF, perf, ftrace).
  • Understanding of deep learning inference workloads and their hardware execution characteristics.
  • Experience with profiling and performance tuning of system software on accelerator or SoC platforms.

Why Join FuriosaAI

  • The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.
  • With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.

Additional Information

  • recruit@furiosa.ai
Role:
Senior Software Engineer, Runtime

Company profile

FuriosaAI

furiosa.ai

FuriosaAI is a South Korean chip company that designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, generative AI, LLMs and other demanding workloads. Its mission is to build AI chips that efficiently run advanced AI models and make AI computing more sustainable. Its products include the Gen 1 Vision NPU and the RNGD data center accelerator, supported by the Furiosa SDK. The company has its Korea HQ in Seoul and offices in Santa Clara, Munich, Lisbon and Singapore.

More jobs at FuriosaAI

Similar Software Engineer jobs at other companies