GPU / Accelerator software · due diligence dossier

不是职位清单,
逐岗审查

只发布已经打开官方详情页、读完整 JD,并验证 Apply 入口或完整申请表的岗位。每一份 review 都明确区分岗位级别、你的已验证证据、缺口和是否值得现在投。

01 · SOURCE公司官网或第一方 ATS,不把搜索摘要当作在岗证明。
02 · STATUS完整 JD 可读,并能进入 Apply 或加载实际申请表。
03 · REVIEW拆开职责、硬要求、加分项、地点、薪酬和级别证据。
04 · FIT只用已验证的候选人证据;缺什么就明确写什么。
Decision
Company
Region 显示 12 / 12

12 份完整 review

排序按当前行动价值,而不是公司名气。岗位的 market level 与你的 candidate fit 分开判断:open-level 不等于 strong match。

CerebrasOfficial + Apply verifiedTrue new grad

Software Engineer - New Grad 2026

当前最值得投。这是唯一同时满足 Toronto 主地点、2026 New Grad、C/C++ 系统方向和可用申请表的岗位。缺口存在,但官方允许用课程、项目和接触经验证明潜力。

Location
官方列表:Toronto, CAN;JD 另写 Toronto 或 Sunnyvale · Hybrid
Employment
Full time · 2026 / recent graduate
Official ID
Ashby 99c289fa…
Posted
2026-05-11
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 设计、实现、测试和调试影响性能与可靠性的系统软件。
  • 参与靠近硬件与网络的低层组件、系统 bring-up 和性能优化。
  • 建设 observability、reliability、scalability 工具,并跨 hardware / firmware / compiler / infrastructure 协作。

Hard + preferred

  • 相关专业、2026 或近期毕业;熟练 C/C++。
  • 对 sockets/networking、embedded、OS、driver、distributed 或 network performance 有兴趣或接触。
  • 加分:Linux/debugger、TCP/RDMA/RPC/Wireshark、driver/embedded/distributed、性能或并发。

Verified overlap

  • UofT CS,预计 2026 毕业,时间与级别精确匹配。
  • 已验证 C/C++、Linux/UNIX、POSIX、pthreads、syscalls、gdb 课程证据。
  • 已有测试、API、协作与技术沟通的工作证据。

Evidence gaps

  • 没有 sockets / TCP / RDMA / RPC 实作。
  • 没有 device-driver 或 embedded 实现。
  • 没有 distributed-systems artifact;debugger 证据仍偏课程级。

Portability + action

T1 · 81。C/C++ 系统、硬件侧调试、bring-up、observability 和性能/可靠性方法可迁移到 AMD/NVIDIA。建议现在做定制简历并投递,不等待所有加分项补齐。

Official proof

Ashby board 数据显示 isListed=true;职位页和申请页均实际渲染,Name、Email、Resume、出口管制问题及 Submit Application 可见。

打开官方职位 ↗
AMDOfficial + Apply verifiedLevel unknown · stretch

Profiling Tool Development Engineer

三份 AMD 中基础重合最好。你的 C++/Python/Linux/Docker/API/testing 能支撑投递叙事,但 production profiling、async networking 与 multi-node data transfer 仍是核心缺口。

Location
Markham, Canada · Hybrid
Comp
CAD 132,000–198,000 / year
Req
80186
Constraint
Not eligible for visa sponsorship
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 为 ROCm profiling stack 设计、编码、测试、集成功能与修复。
  • 建立数据处理抽象,提升 modularity 与 interoperability。
  • 支持 multi-node distributed system 的 scalable data transfer 与 production QA。

Hard + preferred

  • 明确要求快速理解/编写复杂代码并能解释技术概念;相关 BS/MS 或等效背景。
  • Preferred:advanced C++/Python、API、大型应用集成。
  • Preferred:multithreaded async、networking、client/server、distributed、Linux/GitHub/Docker、performance methods。

Verified overlap

  • C++、Python、Linux、GitHub、Docker、数据库均有已验证证据。
  • 有 REST/API 集成、测试与跨团队沟通经验。
  • 系统课程覆盖 POSIX/pthreads,能支撑基础并发叙事。

Evidence gaps

  • 没有 advanced production C++ 的强证明。
  • 没有 ROCm profiler/tool contribution 或 profiler artifact。
  • 没有 async networking、distributed model、multi-node transfer 实现。

Action

选择性投递。职位没有公布 experience band,不能标 entry-level;建议简历聚焦复杂系统代码、testing、API 与 Linux,不夸大 ROCm/driver 经验。

Official proof

AMD Careers 当前渲染完整 JD,明确写明 existing vacancy;Apply 链接可进入 AMD iCIMS 申请入口。

打开官方职位 ↗
TenstorrentBoard + form verifiedOpen-level · stretch

Software Engineer, Metal Runtime (API & Abstractions)

最值得投的 Tenstorrent 岗位。底层 C/C++、host/device abstraction、memory movement 与性能闭环高度可迁移;但 open-level 不是 new grad,你仍缺 accelerator runtime/API 的直接作品。

Location
Toronto / Austin / Santa Clara · Hybrid
Level
Various experience levels; interview-calibrated
Job ID
Greenhouse 5192671007
Comp note
$100k–$500k company-wide target incl. variable; not a Toronto-specific band
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 设计 host/device runtime API 与硬件能力抽象。
  • 实现、调试 runtime,并与 kernel/compiler 团队完成 integration。
  • 优化 performance、API usability、maintainability 与长期接口质量。

Qualification signals

  • 强 C/C++,靠近硬件的 low-level systems。
  • 理解 concurrency、processors、memory movement、performance-critical systems。
  • 有 API/abstraction design 意识;JD 没有独立 preferred section,也没有最低学历/年限。

Verified overlap

  • C/C++、Linux/POSIX、pthreads 与 API/testing 是真实基础。
  • 概念层 GPU architecture/ROCm 能帮助理解 host/device 边界。
  • 已有工程协作和调试能力证据。

Evidence gaps

  • 没有 accelerator runtime 或 host/device API 实现。
  • 没有 profiler-backed performance artifact。
  • 没有 CUDA/HIP 实作或亲自完成的并发 benchmark。

Portability + action

T1 · 82,约 30% 厂商专有,切换成本中等。建议选择性投递,同时优先补一个可运行的 GPU kernel/runtime benchmark。

Official proof

官方 Greenhouse detail page/API 返回完整 posting;ID 位于当前 board-wide listing,页面显示完整申请表与 Submit application。该数字是 Greenhouse job/posting ID,不是公司标注的 requisition。

打开官方职位 ↗
CerebrasOfficial + form verifiedTrue new grad

Kernel Engineer - New Grad

技术迁移性极高,但属于美国地点 stretch。岗位真正接受课程、实习、研究或个人项目作为调试证据;你的基础对口,但 kernel、CUDA/OpenCL/DSL 与 profiler 证据尚未建立。

Location
Sunnyvale, CA · Hybrid · US-only
Employment
Full time · New Grad
Official ID
Ashby 9c7da4b8…
Posted
2026-07-23
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 为 WSE 实现和优化 ML、linear algebra 与 HPC kernels。
  • 用 Cerebras Software Language 将并行算法映射到定制架构。
  • 用数学分析、profiling、tests 调查 correctness、performance 与 utilization。

Hard + preferred

  • 相关 BS/MS/PhD;扎实 C++,熟悉 Python;architecture、DSA、debugging 基础。
  • 关注 low-level、parallel、performance 或 HW/SW co-design。
  • 加分:kernel/compiler/HPC、multithreading、accelerator、assembly/CUDA/OpenCL/DSL、ML framework、linear algebra、profiling。

Verified overlap

  • C++、Python、DSA/系统课程与 gdb 调试证据。
  • 具备 GPU architecture 与 memory hierarchy 的概念基础。
  • 学历/毕业阶段符合 New Grad 画像。

Evidence gaps

  • 没有 accelerator-kernel 实现。
  • 没有 CUDA/OpenCL/DSL 执行记录。
  • 没有 numerical/linear-algebra kernel 或 profiler-backed GPU benchmark。

Portability + action

T0 · 88,约 20% 厂商专有;kernel/performance 方法高度可迁移到 CUDA/HIP。先确认美国工作授权/relocation 条件,再选择性投递;页面没有为你确认 sponsorship。

Official proof

Ashby board 当前列出该岗位;申请页实际加载 Resume、Transcript、Education、relocation、出口管制问题及 Submit Application。

打开官方职位 ↗
QualcommOfficial + Apply verifiedOpen-level · stretch

Machine Learning Framework, Compiler & Performance Engineer

之前的“未核验”结论已纠正。官方 Qualcomm Careers 页面和 Apply Now 都真实存在。Markham + 无最低年限使它值得选择性投递,但 compiler、PyTorch/ONNX 和 SoC simulation 仍是实质缺口。

Location
Markham, Ontario · Flexible Location(不可写成 remote)
Comp
$99,500–$149,300;官方文本未标币种
IDs
446720263662 · Job ID 3094691
Posted
2026-08-04 · Replacement Position
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 开发 production / exploratory ML-AI compilers 与 workload compilation algorithms。
  • 连接 PyTorch 与 Qualcomm compiler flows,分析 performance/area/power trade-off。
  • 用 C++/Python 构建 SoC simulation,做 pre-silicon performance prediction 和 bottleneck debug。

Hard + preferred

  • Minimum:相关 Bachelor's;没有写工作年限。
  • 核心强调 algorithm、performance analysis、debugging、analysis 与 communication。
  • 加分:强 C++/Python/OOD、compiler、PyTorch/ONNX、Git/Jenkins/Docker/CI、on-silicon debug、architecture/digital circuits/simulator。

Verified overlap

  • C++、Python、Docker、GitHub、testing 与 API integration 有证据。
  • 有 AI provider integration 与 LoRA evaluation 的相邻经验。
  • 有 architecture/GPU 概念基础与跨团队沟通证据。

Evidence gaps

  • 没有 compiler implementation 或 lowering/codegen。
  • 没有 PyTorch/ONNX framework integration。
  • 没有 SoC simulation、pre-silicon prediction、on-silicon debug 或 event-driven simulator。

Portability + action

T0 · 86,估计约 25% 厂商专有;compiler/framework/architecture/performance 方法可迁移。建议选择性投递,但简历必须把 AI integration 与 compiler engineering 清楚分开。

Official proof

Qualcomm Careers 完整职位页可见,Apply Now 指向官方申请入口并进入登录/创建账号流程。官方数据同时出现 onsite 与 flexible-within-country,不能解读为 remote。

打开官方职位 ↗
AMDOfficial + Apply verifiedHigh stretch

Platform Emulation Software Engineer - DC GPU

方向对,但证据差距很大。你有系统与 GPU 架构基础;岗位却要求真实 CUDA/HIP workload、Tcl、pre-silicon emulator、coherence 与 boot-flow 能力。可作为高价值 stretch,不应包装成“已匹配”。

Location
Markham preferred;Austin also open · Hybrid
Comp
CAD 129,600–194,400 / year
Req
89890
Constraint
Not eligible for visa sponsorship
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 开发 CUDA/C++ 与 CUDA/HIP GPU apps,用于 pre-silicon verification。
  • 构建 emulation debug tooling、automation、benchmarks 与 data analysis。
  • 运行 AI/ML functional/performance workloads,定位 hardware/performance failures。

Requirements — corrected

  • 强 GPU/CPU architecture、memory hierarchy、interconnect、cache;excellent C/C++、Python、Tcl。
  • shared-memory concurrency、relaxed memory models、cache coherence、Linux/Unix shell、software architecture/boot flow 都是要求项。
  • 加分才是 SystemVerilog/waveform、ML profiling、ROCm/OpenCL/OpenGL/Vulkan。

Verified overlap

  • C/C++、Python、Linux shell、POSIX/pthreads。
  • GPU architecture、AMDGPU、ROCm 属于概念层基础。
  • 测试、debugging 与沟通能力可作为辅助证据。

Evidence gaps

  • 没有执行过的 CUDA/HIP application。
  • 没有 Tcl、emulator、cache-coherence implementation。
  • 没有 Verilog/waveform debug 或 profiler-backed benchmark。

Action

只做选择性 stretch 投递。若时间有限,优先 Cerebras Toronto、AMD 80186 和 TT Metal。这个 JD 更适合作为下一阶段 evidence checklist。

Official proof

AMD Careers 当前显示完整 JD、existing vacancy 声明与 Apply;官方 Apply 可达加拿大 iCIMS。要求与 preferred 的段落归属已逐项核对。

打开官方职位 ↗
AMDOfficial + Apply verifiedWatch · stretch

AI Model, Framework, and GPU Software Engineer – Agentic AI

现在不优先。你有 C/C++/Python/Linux 与 AI integration 邻接经验,但岗位真正拥有的是 model bring-up、framework/inference stack 和 GPU operator performance;这些尚未被你的证据覆盖。

Location
Markham, Canada · Hybrid
Comp
CAD 132,000–198,000 / year
Req
88870
Level
No years range / no new-grad wording
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 让模型与 frameworks 在 AMD GPUs 上高效执行。
  • 支持 framework/inference/agentic stack、model bring-up 与 performance analysis。
  • 从模型架构到 GPU acceleration 做 cross-stack optimization,并参与 OSS/customer enablement。

Preferred profile

  • C/C++ and/or Python on Linux/Windows。
  • AI/ML/inference;PyTorch、vLLM、llama.cpp。
  • GPU acceleration/performance/compiler/low-level、operator performance、open source;相关学位 preferred。

Verified overlap

  • C/C++、Python、Linux 与跨团队软件开发。
  • ImplicitCAD 提供 AI provider integration 与 model-evaluation 邻接证据。
  • 具备测试与开源工程工作方式基础。

Evidence gaps

  • 没有 PyTorch/vLLM/llama.cpp development。
  • 没有 GPU execution 或 operator optimization artifact。
  • 没有 compiler/MLIR/Triton 证据。

Action

先补证据再投。AI API/LoRA evaluation 不能代替 framework internals 或 GPU performance engineering。最小下一步是可复现的 GPU workload + profiler 优化报告。

Official proof

AMD Careers 当前显示完整 JD、existing vacancy 声明和有效 Apply 链接;标题没有 Senior/Staff,但也没有 entry/open-level 证据。

打开官方职位 ↗
TenstorrentOfficial + form verifiedOpen-level · high stretch

Software Engineer, Acceleration Kernel Development

路线价值高于当前投递价值。这是最直接的 kernel/performance 能力模板:low-level C/C++、cycles、memory、bandwidth、profiling。也正因为如此,你目前缺失的证据几乎就是岗位核心。

Location
Toronto, Ontario · Hybrid
Level
Various experience levels
Job ID
Greenhouse 4155609007
Comp note
$100k–$500k company-wide target incl. variable; not Toronto-specific
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 为 ML workloads 写 low-level code 与高性能 parallel algorithms。
  • 针对 cycles、instruction latency、memory 和 bandwidth 做 kernel tuning。
  • 负责 debugging、profiling、integration 与长期维护。

Qualification signals

  • 强 C/C++。
  • ML/HPC compute-kernel optimization。
  • instruction-level、latency、memory、bandwidth reasoning;debug/profile ownership。

Verified overlap

  • C/C++、systems、DSA 与基础性能意识。
  • 概念层 GPU architecture / memory hierarchy。
  • 测试与调试工作方式。

Evidence gaps

  • 没有 compute-kernel implementation。
  • 没有 accelerator performance measurements。
  • 没有 instruction/bandwidth tuning 或 profiler-backed artifact。

Portability + action

T0 · 88,约 25% 厂商专有,切换成本中等。把它当作 GPU reduction/scan 项目的验收标准;完成 artifact 后再升级为投递目标。

Official proof

官方 Greenhouse 职位页完整渲染,当前存在申请表与 Submit application。页面明确接受不同经验级别,但不等于明确 New Grad。

打开官方职位 ↗
TenstorrentBoard + form verifiedOpen-level · high stretch

Software Engineer, TT-Distributed

真实开放,但暂不建议优先投。它要求的是 accelerator cluster 的 distributed execution:MPI-like technology、IPC/socket、fault tolerance、multi-node profiling。当前证据没有覆盖这个闭环。

Location
结构化字段:Toronto;JD 正文:Santa Clara / Austin / Toronto(官方内部不一致)
Level
Various experience levels
Job ID
Greenhouse 4711506007
Comp note
$100k–$500k company-wide target incl. variable; not Toronto-specific
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 构建 accelerator/CPU cluster distributed systems。
  • 设计 data-parallel、tensor-parallel API 与 communication primitives。
  • 建设 testing、debugging、profiling、monitoring 和 cluster bring-up 工具。

Qualification signals

  • C/C++、systems/OS/distributed foundations。
  • IPC、sockets、cluster coordination、scalability、fault tolerance、multi-node performance。
  • MPI 出现在职责/工具信号中,不是官方标注的 preferred qualification。

Verified overlap

  • C/C++、Linux/POSIX 与系统课程基础。
  • 测试、调试和数据处理经验。
  • 有并发概念与 pthreads 证据,但范围有限。

Evidence gaps

  • 没有 MPI/collectives。
  • 没有 socket/IPC、fault-tolerant distributed implementation。
  • 没有 multi-node profiling、monitoring 或 cluster bring-up。

Portability + action

T1 · 73,约 35% 厂商专有,切换成本中高。保留为学习信号;先完成 GPU artifact,再考虑 bounded MPI/collectives branch。

Official proof

Greenhouse detail/API 均 200,ID 在当前 board listing,申请表完整可提交。位置字段冲突在这里原样披露,没有静默合并。

打开官方职位 ↗
TenstorrentBoard + form verifiedOpen-level · high stretch

Software Engineer, TT-Fabric

技术门槛高,适合作为路线目标。岗位核心是 low-level fabric/networking library、protocol、synchronization 与 data movement,而不是一般 C++ 应用开发。

Location
Toronto / Austin / Santa Clara · Hybrid
Level
Various experience levels
Job ID
Greenhouse 4645584007
Comp note
$100k–$500k company-wide target incl. variable; not Toronto-specific
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 设计与维护 TT-Fabric low-level networking library。
  • 协调大量 AI processors,优化 protocols、synchronization 与 data movement。
  • 把 Fabric API 接入 programming model,并参与长期 distributed-stack architecture。

Qualification signals

  • 深度 C/C++。
  • low-level 或 bare-metal;hardware/software interaction。
  • protocol-level tuning、networking、synchronization、cluster communication。

Verified overlap

  • C/C++、Linux/POSIX、pthreads 与系统基础。
  • GPU/architecture 概念知识有助于 HW/SW reasoning。
  • 测试与 debugging 工作方式。

Evidence gaps

  • 没有 bare-metal implementation。
  • 没有 network protocol、RDMA/interconnect library。
  • 没有 distributed synchronization/collectives 或低延迟 data-movement benchmark。

Portability + action

T1 · 72,约 35% 厂商专有,切换成本中高。知识可映射到 RDMA、NCCL/RCCL 与 GPU interconnect,但当前只建议观察和学习。

Official proof

Greenhouse 页面/API、当前全量 board listing 与 Submit application 表单全部存在;未发现最低学历或年限。

打开官方职位 ↗
TenstorrentOfficial + form verifiedOpen-level · high stretch

Systems Engineer, Data Center Debug

它是硬件 bring-up 岗,不是普通 systems role。需要跨 chip/system/firmware/software 的 post-silicon root cause,且涉及实验室设备、PCIe、SerDes、GDDR、power/thermal。

Location
Toronto, Ontario · Hybrid
Level
Various experience levels
Job ID
Greenhouse 5143663007
Comp
No compensation text captured on this posting
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 在 chips、systems、firmware、software 之间做 post-silicon debug。
  • 处理 power/thermal、GDDR、PCIe、Ethernet/SerDes、NoC failure。
  • 用 C/Python 构建 diagnostics,读取 schematics/registers/telemetry,并操作 lab equipment。

Qualification signals

  • hardware debug、processor architecture、trace methodology。
  • 跨层 root-cause analysis 与 automation。
  • firmware、memory、PCIe/networking、lab debug 工具。

Verified overlap

  • C、Python、Linux 与 debugging/testing 基础。
  • architecture、AMDGPU/ROCm 概念知识。
  • 自动化与跨团队沟通经验。

Evidence gaps

  • 没有 post-silicon bring-up。
  • 没有 board/lab equipment、PCIe/SerDes/GDDR failure analysis。
  • 没有 hardware-facing diagnostic-tool implementation。

Portability + action

T1 · 76,约 30% 厂商专有,切换成本中等。诊断方法可迁移,但硬件接口与 lab 能力是硬门槛;现在只作为市场信号。

Official proof

官方 Greenhouse 页面完整渲染,并存在实际申请表和 Submit application。Open-level 只说明面试可定级,不降低 bring-up 技术门槛。

打开官方职位 ↗
TenstorrentOfficial + form verifiedOpen-level · very high stretch

Software Engineer, AI Compiler

可迁移性最高,当前匹配度最低。JD 虽写 various experience levels,但同时包含 lead、roadmap、mentorship 语义;没有 MLIR/LLVM 或 compiler 实现时,不应把它当普通 early-career opening。

Location
Austin, Texas · Hybrid · US-only
Level
Various levels; high ownership wording
Job ID
Greenhouse 4717175007
Comp note
$100k–$500k company-wide target incl. variable
完整审查 · 职责 / 要求 / 证据 / 缺口 / 迁移性 / 核验链

Role owns

  • 构建 AI compiler infrastructure、dialects 与 optimization passes。
  • 连接 framework execution、graph lowering、runtime 与 Tenstorrent backend。
  • 承担 cross-team tooling、roadmap、technical leadership/mentorship。

Qualification signals

  • compiler 或类似 systems experience。
  • 强 C++/Python;MLIR/LLVM、dialect/pass。
  • 理解 PyTorch/TensorFlow/JAX execution 与可靠 cross-team delivery。

Verified overlap

  • C++、Python、DSA 与系统基础。
  • AI integration 提供 framework 使用层的邻接经验。
  • 测试、工具和跨团队交付方式可迁移。

Evidence gaps

  • 没有 compiler implementation。
  • 没有 LLVM/MLIR dialect/pass。
  • 没有 lowering/codegen 或 deep-learning framework internals。

Portability + action

T0 · 90,约 25% 厂商专有,切换成本低到中等。MLIR/LLVM 概念高度通用,但这是未来路线目标,不是当前投递建议。

Official proof

官方 Greenhouse 页面和申请表当前可用,Submit application 存在。地点只列 Austin;页面级别文案与职责深度的张力已显式披露。

打开官方职位 ↗

没有发布的职位

“被搜到过”不等于“现在真实开放”。下面这些 lead 被主动挡在主列表之外,避免旧缓存和失效 URL 伪装成机会。

AMD 90131上午曾显示 NPU Diagnostics Software Engineer;15:46 EDT 官方详情页提示职位不可用,iCIMS 跳转含 notFound=1Removed
NVIDIA JR2016937搜索缓存指向 Toronto New Grad compiler lead,但官方 Workday 当前显示页面不存在。Not published
NVIDIA JR2014274搜索缓存指向 ASIC methodology lead,但官方 Workday 当前显示页面不存在。Not published