High-Performance Rendering Framework on Stream Architectures
-
Updated
Oct 11, 2026 - C++
High-Performance Rendering Framework on Stream Architectures
AscendMate · 昇腾部署易用一指禅 —— 面向昇腾智算服务器的环境搭建、模型微调、推理部署、算子开发手册
Ascend Assistant · 昇腾服务器助手 —— Agent Skill,帮助用 AI 操作/查询/排障昇腾智算服务器(环境检测/排障/命令生成/部署引导/性能调优),与 AscendMate 手册深度联动
Ascend NPU fork of nanochat for LLM training with torch_npu/HCCL (experimental)
End-to-end edge AI scene recognition on Huawei Ascend 310: PyTorch, ONNX, C++ inference, V4L2, and MJPEG streaming.
NPU MFU Analyzer - 大模型训练性能分析工具 | LLM Training Performance Analyzer for Ascend NPU
Native AscendC Mamba2 selective scan / SSD forward-backward custom operator for Huawei Ascend 910B3 and 950PR, with CANN, torch_npu, A100 benchmarks and msprof profiling.
Disaggregated Prefill-Decode serving engine connecting NVIDIA Blackwell (SM120) and Huawei Ascend 910B2 (CANN) across physical nodes.
High-performance FFT operator library for Huawei Ascend, applying cuButterfly's parameterized hybrid-dataflow method with hardware profiling, performance modeling, and autotuning.
Industrial defect detection on Huawei Ascend NPU with YOLO, ONNX-to-OM conversion, and C/C++/Python inference.
驱动只有网页终端、在WAF后的竞赛/云pod:反向SSH隧道+命令通道,从自己机器全自动发命令拿结果(含踩坑与安全加固)
Deploy GLM-4.6V-Flash (9B dense VLM) on Huawei Ascend 910B NPU with vLLM - multimodal, OpenAI API, single/dual-card serving, reproducible benchmarks.
To associate your repository with the huawei-ascend topic, visit your repo's landing page and select "manage topics."