Job description
The Private Cloud Compute (PCC) team builds Apple's privacy-preserving cloud inference platform — the system that lets generative-AI features on iPhone, iPad, Mac, and Vision Pro reach into the cloud without compromising user privacy. We are looking for a Software Engineer to join the team that builds the on-device client side of that pipeline: the system frameworks and sandboxed services that every Apple Intelligence and Foundation Models request flows through before it ever leaves the device.
This is system-level Apple platform work. You will design and ship the Swift frameworks that turn high-level inference requests into a compact wire protocol, manage long-lived connections, stream multimodal output back to client apps - all under strict process isolation and privacy guarantees. Your code will run on every device that uses Apple Intelligence — and the bar for correctness, performance, and privacy is set accordingly. You will be a key contributor to the Swift frameworks and XPC services that connect on-device apps and system daemons to PCC. Day-to-day responsibilities include:
Designing and evolving the on-device client SDKs and the XPC services they sit behind, balancing API ergonomics for client teams with the strict process-isolation and attestation guarantees required by PCC.
Implementing high-throughput, low-latency request paths: protobuf serialization of inference payloads, payload compression, multi-turn session state and lifecycle, retry/cancellation semantics, and streaming of incremental model output back to clients.
Owning concurrency and resource management across async/await, AsyncStream, and GCD — including correct cancellation, back-pressure, and lifetime management of long-lived XPC connections.
Building the telemetry, signposts, and privacy-audited logging that let on-call engineers diagnose request failures across device, cloud, and inference layers without exposing user data.
Collaborating closely with the PCC server-side, Foundation Models, attestation, and platform-frameworks teams; writing design documents, leading reviews, and mentoring engineers on Apple platform fundamentals.
Profiling on real hardware with Instruments and driving down memory, latency, and energy costs in the request path — these libraries run on every device, on every inference, so every microsecond matters.
This is system-level Apple platform work. You will design and ship the Swift frameworks that turn high-level inference requests into a compact wire protocol, manage long-lived connections, stream multimodal output back to client apps - all under strict process isolation and privacy guarantees. Your code will run on every device that uses Apple Intelligence — and the bar for correctness, performance, and privacy is set accordingly. You will be a key contributor to the Swift frameworks and XPC services that connect on-device apps and system daemons to PCC. Day-to-day responsibilities include:
Designing and evolving the on-device client SDKs and the XPC services they sit behind, balancing API ergonomics for client teams with the strict process-isolation and attestation guarantees required by PCC.
Implementing high-throughput, low-latency request paths: protobuf serialization of inference payloads, payload compression, multi-turn session state and lifecycle, retry/cancellation semantics, and streaming of incremental model output back to clients.
Owning concurrency and resource management across async/await, AsyncStream, and GCD — including correct cancellation, back-pressure, and lifetime management of long-lived XPC connections.
Building the telemetry, signposts, and privacy-audited logging that let on-call engineers diagnose request failures across device, cloud, and inference layers without exposing user data.
Collaborating closely with the PCC server-side, Foundation Models, attestation, and platform-frameworks teams; writing design documents, leading reviews, and mentoring engineers on Apple platform fundamentals.
Profiling on real hardware with Instruments and driving down memory, latency, and energy costs in the request path — these libraries run on every device, on every inference, so every microsecond matters.