Skip to content

[QNN EP] Add use_native_matmul session config for native MatMul lowering - #821

Open
qti-mbadnara wants to merge 10 commits into
mainfrom
dev/qti-mbadnara/poc-disable-matmul-to-fc
Open

qti-mbadnara wants to merge 10 commits into
mainfrom
dev/qti-mbadnara/poc-disable-matmul-to-fc

Conversation

@qti-mbadnara

@qti-mbadnara qti-mbadnara commented Sep 12, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Adds an ep.QNN.use_native_matmul session config key (default "0") that routes MatMul and Gemm operators through QNN_OP_MAT_MUL instead of QNN_OP_FULLY_CONNECTED, eliminating the host-side weight transposition that occurs during session creation.


Motivation & Context

ORT QNN EP lowers MatMul (with a rank-2 static weight) and all Gemm operators to QNN_OP_FULLY_CONNECTED. FC requires the weight in [N, K] layout, but ONNX stores it in [K, N], so the builder pre-transposes every static weight at session creation time.

QNN_OP_MAT_MUL accepts weight in the natural ONNX [K, N] orientation, making it a drop-in replacement that avoids this transposition entirely. Compared to QNN_OP_FULLY_CONNECTED, it also supports a larger max rank (5 vs. 4), which removes the need to flatten leading activation dimensions for batched inputs.

Paths left on FC (not affected by this flag):

  • Block-quantized (BW_FLOAT_BLOCK) Gemm — BQ weight layout and quantization params are FC-specific
  • Absorbed-Reshape QDQ Gemm (output_reshape_node != nullptr) — narrow QDQ pattern, left for follow-up
  • LPBQ MatMul / LPBQ Gemm fusions — unaffected; LPBQ MatMul already emits QNN_OP_MAT_MUL natively

Supported Configuration

Config key Value Behavior
ep.QNN.use_native_matmul "0" (default) MatMul and Gemm route to QNN_OP_FULLY_CONNECTED. Weight transposed to [N, K] at session creation.
ep.QNN.use_native_matmul "1" MatMul and Gemm route to QNN_OP_MAT_MUL. Weight stays in [K, N]; transpose_in0/transpose_in1 params set from transA/transB attributes.

Unsupported Configuration

  • Block-quantized (BW_FLOAT_BLOCK) Gemm — always emits QNN_OP_FULLY_CONNECTED regardless of this flag; BQ weight encoding is FC-specific.
  • Absorbed-Reshape QDQ Gemm (output_reshape_node != nullptr) — left on FC; targeted for follow-up.
  • LPBQ MatMul / LPBQ Gemm fusions — independent node group fusions, unaffected.

Data Types

Both QNN_OP_FULLY_CONNECTED and QNN_OP_MAT_MUL share identical data type support. Switching between them has no effect on supported dtypes.

Backend Supported dtypes
CPU FP32
HTP FP16, BF16, INT8, INT16 (Firewheel only), BW_FLOAT_BLOCK FP16+SFIXED8 (V69+)
GPU FP16, FP32

Files Changed

File Change
onnxruntime/core/providers/qnn/builder/qnn_model_wrapper.h Added bool use_native_matmul = false to ModelSettings. Threaded to op builders via GetModelSettings().
onnxruntime/core/providers/qnn/qnn_execution_provider.cc Parses ep.QNN.use_native_matmul via ParseBoolOption (default false).
onnxruntime/core/providers/qnn/builder/opbuilder/matmul_op_builder.cc Guard in CheckInputs: forces use_fully_connected = false when flag is set. Existing QNN_OP_MAT_MUL path handles rank-2 static weights correctly without transposition.
onnxruntime/core/providers/qnn/builder/opbuilder/qlinear_matmul_op_builder.cc Same guard in DecideUseFullyConnected.
onnxruntime/core/providers/qnn/builder/qnn_node_group/reshape_gemm_fusion.cc Early return nullptr in TryFusion2/3/4 when flag is set, allowing nodes to fall through to GemmOpBuilder.
onnxruntime/core/providers/qnn/builder/opbuilder/gemm_op_builder.cc (a) ProcessInputs: skips host-side weight transposition when flag is set. (b) ProcessAttributesAndOutputs: normal path and FC+Add decomposition path emit QNN_OP_MAT_MUL with transpose params derived from transA/transB.
docs/execution_providers/QNN-ExecutionProvider.md Documents the new session config key.
onnxruntime/test/providers/qnn/matmul_test.cc 4 new tests: MatMulDisableFC_2D_StaticWeight, MatMulDisableFC_3D_StaticWeight, MatMulDefaultUsesFC, MatMulOptInUsesMatMul.
onnxruntime/test/providers/qnn/gemm_test.cc 3 new tests: GemmDisableFC_NoBias_TransB0, GemmDisableFC_NoBias_TransB1, GemmDisableFC_WithBias.

@qti-mbadnara
qti-mbadnara enabled auto-merge (squash) September 24, 2026 02:24
Comment thread onnxruntime/core/providers/qnn/builder/qnn_node_group/reshape_gemm_fusion.cc Outdated
Comment thread onnxruntime/core/providers/qnn/builder/qnn_node_group/reshape_gemm_fusion.cc Outdated
Comment thread onnxruntime/core/providers/qnn/builder/opbuilder/qlinear_matmul_op_builder.cc Outdated
Comment thread onnxruntime/core/providers/qnn/builder/opbuilder/matmul_op_builder.cc Outdated
Comment thread onnxruntime/core/providers/qnn/builder/opbuilder/matmul_op_builder.cc Outdated
Comment thread onnxruntime/core/providers/qnn/builder/opbuilder/qlinear_matmul_op_builder.cc Outdated
Comment thread onnxruntime/core/providers/qnn/builder/qnn_node_group/reshape_gemm_fusion.cc Outdated
Comment thread onnxruntime/core/providers/qnn/qnn_execution_provider.cc Outdated
Comment thread docs/execution_providers/QNN-ExecutionProvider.md Outdated
@qti-mbadnara qti-mbadnara changed the title [QNN EP] Add disable_matmul_to_fc session config for native MatMul lowering [QNN EP] Add use_native_matmul session config for native MatMul lowering Sep 29, 2026
Signed-off-by: qti-mbadnara <mbadnara@qti.qualcomm.com>
Comment thread docs/execution_providers/QNN-ExecutionProvider.md
@qti-mbadnara
qti-mbadnara force-pushed the dev/qti-mbadnara/poc-disable-matmul-to-fc branch 2 times, most recently from 547d6d9 to 8ab0ebf Compare September 30, 2026 01:17

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants