[QNN EP] Add use_native_matmul session config for native MatMul lowering - #821
Open
qti-mbadnara wants to merge 10 commits into
Open
qti-mbadnara wants to merge 10 commits into
qti-mbadnara wants to merge 10 commits into
Conversation
qti-mbadnara
requested review from
minfhong-qti,
qti-ashwshan,
qti-chuteng,
qti-kromero,
qti-shubham,
qti-yuduo,
quic-calvnguy,
tirupath-qti and
yuhuchua-qti
as code owners
September 12, 2026 03:39
Signed-off-by: qti-mbadnara <mbadnara@qti.qualcomm.com>
qti-mbadnara
enabled auto-merge (squash)
September 24, 2026 02:24
qti-kromero
reviewed
Sep 28, 2026
Signed-off-by: qti-mbadnara <mbadnara@qti.qualcomm.com>
qti-kromero
approved these changes
Sep 29, 2026
qti-mbadnara
force-pushed
the
dev/qti-mbadnara/poc-disable-matmul-to-fc
branch
2 times, most recently
from
September 30, 2026 01:17
547d6d9 to
8ab0ebf
Compare
Signed-off-by: qti-mbadnara <mbadnara@qti.qualcomm.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Adds an
ep.QNN.use_native_matmulsession config key (default"0") that routesMatMulandGemmoperators throughQNN_OP_MAT_MULinstead ofQNN_OP_FULLY_CONNECTED, eliminating the host-side weight transposition that occurs during session creation.Motivation & Context
ORT QNN EP lowers
MatMul(with a rank-2 static weight) and allGemmoperators toQNN_OP_FULLY_CONNECTED. FC requires the weight in[N, K]layout, but ONNX stores it in[K, N], so the builder pre-transposes every static weight at session creation time.QNN_OP_MAT_MULaccepts weight in the natural ONNX[K, N]orientation, making it a drop-in replacement that avoids this transposition entirely. Compared toQNN_OP_FULLY_CONNECTED, it also supports a larger max rank (5 vs. 4), which removes the need to flatten leading activation dimensions for batched inputs.Paths left on FC (not affected by this flag):
output_reshape_node != nullptr) — narrow QDQ pattern, left for follow-upQNN_OP_MAT_MULnativelySupported Configuration
ep.QNN.use_native_matmul"0"(default)QNN_OP_FULLY_CONNECTED. Weight transposed to[N, K]at session creation.ep.QNN.use_native_matmul"1"QNN_OP_MAT_MUL. Weight stays in[K, N];transpose_in0/transpose_in1params set fromtransA/transBattributes.Unsupported Configuration
QNN_OP_FULLY_CONNECTEDregardless of this flag; BQ weight encoding is FC-specific.output_reshape_node != nullptr) — left on FC; targeted for follow-up.Data Types
Both
QNN_OP_FULLY_CONNECTEDandQNN_OP_MAT_MULshare identical data type support. Switching between them has no effect on supported dtypes.Files Changed
onnxruntime/core/providers/qnn/builder/qnn_model_wrapper.hbool use_native_matmul = falsetoModelSettings. Threaded to op builders viaGetModelSettings().onnxruntime/core/providers/qnn/qnn_execution_provider.ccep.QNN.use_native_matmulviaParseBoolOption(defaultfalse).onnxruntime/core/providers/qnn/builder/opbuilder/matmul_op_builder.ccCheckInputs: forcesuse_fully_connected = falsewhen flag is set. ExistingQNN_OP_MAT_MULpath handles rank-2 static weights correctly without transposition.onnxruntime/core/providers/qnn/builder/opbuilder/qlinear_matmul_op_builder.ccDecideUseFullyConnected.onnxruntime/core/providers/qnn/builder/qnn_node_group/reshape_gemm_fusion.ccreturn nullptrinTryFusion2/3/4when flag is set, allowing nodes to fall through toGemmOpBuilder.onnxruntime/core/providers/qnn/builder/opbuilder/gemm_op_builder.ccProcessInputs: skips host-side weight transposition when flag is set. (b)ProcessAttributesAndOutputs: normal path and FC+Add decomposition path emitQNN_OP_MAT_MULwith transpose params derived fromtransA/transB.docs/execution_providers/QNN-ExecutionProvider.mdonnxruntime/test/providers/qnn/matmul_test.ccMatMulDisableFC_2D_StaticWeight,MatMulDisableFC_3D_StaticWeight,MatMulDefaultUsesFC,MatMulOptInUsesMatMul.onnxruntime/test/providers/qnn/gemm_test.ccGemmDisableFC_NoBias_TransB0,GemmDisableFC_NoBias_TransB1,GemmDisableFC_WithBias.