Skip to content

[QNN EP] Add graph-structure tests for DQQ and CastLoneQ fusions - #539

Open
qti-chuteng wants to merge 2 commits into
mainfrom
dev/chuteng/dqq-castloneq-fusion-tests
Open

qti-chuteng wants to merge 2 commits into
mainfrom
dev/chuteng/dqq-castloneq-fusion-tests

Conversation

@qti-chuteng

Copy link
Copy Markdown
Collaborator

Summary

Adds graph-structure unit tests for the two Q/DQ-to-Convert fusions in the QNN EP — DQQFusion and CastLoneQFusion. Both fusions previously had only accuracy / EP-assignment coverage; nothing asserted that the fused Convert actually appears in the compiled QNN graph. These tests close that gap.

No production code changes — test-only.

Why

Both fusions rewrite a small Q/DQ pattern into a single QNN_OP_CONVERT:

  • DQQFusion: a standalone DequantizeLinear -> QuantizeLinear (same scale type, scalar constant qparams) collapses into one Convert.
  • CastLoneQFusion: a standalone Cast whose parent is not a DequantizeLinear and whose only child is a QuantizeLinear collapses into one Convert.

Standalone Quantize / Dequantize / Cast all have QNN op builders, so whether or not the fusion fires the nodes still land on the QNN EP (ExpectedEPNodeAssignment::All). EP assignment therefore cannot distinguish "fused" from "not fused" — the only reliable signal is the compiled-graph JSON. These tests dump the QNN graph (dump_json_qnn_graph) and assert on the op types via AssertOpInQnnGraph.

Coverage

onnxruntime/test/providers/qnn/qnn_node_group/dq_q_fusion_test.cc

Test Asserts
DQQFusion_U8_Convert uint8 requantize at output → Convert count 1
DQQFusion_U16_Convert uint16 requantize at output → Convert count 1
DQQFusion_ExtraConsumer_NoFusion DQ has a 2nd consumer → no fusion, Convert count 0

onnxruntime/test/providers/qnn/qnn_node_group/cast_lone_q_fusion_test.cc

Test Asserts
CastLoneQFusion_U8_Convert Cast -> Q → Convert count 1, Cast count 0
CastLoneQFusion_NoQChild_NoFusion Cast -> Add (no Q child) → Convert count 0, Cast count 1

Implementation notes

  • Test harness (ResetQnnGraphDir / HasQnnJsonGraph / GetHTPProviderOptions / RunFusionTest) mirrors the existing dql_dq_fusion_test.cc.
  • Model topologies reuse patterns already proven in simple_op_test.cc (BuildDQQConvertAtOutputTestCase, including the scale *= 1.01f nudge so the ORT QDQ optimizer does not fold the DQ -> Q pair away before the QNN EP sees it) and qnn_basic_test.cc (the Cast -> Q -> DQ -> Add topology).
  • Platform guards (#if defined(__aarch64__) || defined(_M_ARM64) || defined(__linux__)) and SKIP_HTP_TEST_ON_ARCH_LESS_THAN_OR_EQUAL_TO(QNN_HTP_DEVICE_ARCH_V68) match sibling fusion tests.
  • New files are picked up automatically by the providers/qnn/qnn_node_group/* CMake glob — no CMake change needed.

Testing

  • lintrunner — clean.
  • These are HTP graph-dump tests gated behind the arm64/linux platform guard; run them with --gtest_filter="QnnHTPBackendTests.DQQFusion*:QnnHTPBackendTests.CastLoneQFusion*" on an HTP-capable environment (arm64 device or x86 HTP emulator).

Add JSON-graph verification tests for the two Q/DQ-to-Convert fusions that
previously had accuracy-only coverage:

- DQQFusion: a lone DequantizeLinear -> QuantizeLinear pair must collapse
  into a single QNN Convert. Covers uint8 and uint16 requantize at the graph
  output, plus a no-fusion case where the DQ has a second consumer.
- CastLoneQFusion: a standalone Cast feeding a QuantizeLinear (with a non-DQ
  parent) must collapse into a single QNN Convert. Covers the uint8 happy
  path plus a no-fusion case where the Cast has no Q child.

Both fusions emit QNN_OP_CONVERT, and standalone Quantize/Dequantize/Cast all
have QNN op builders, so fusion is verified via the compiled-graph JSON
(AssertOpInQnnGraph on "Convert"/"Cast") rather than EP assignment. The test
harness mirrors dql_dq_fusion_test.cc and reuses the model topologies already
proven in simple_op_test.cc and qnn_basic_test.cc.
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.


qti-chuteng seems not to be a GitHub user. You need a GitHub account to be able to sign the CLA. If you have already a GitHub account, please add the email address used for this commit to your account.
You have signed the CLA already but the status is still pending? Let us recheck it.

// The DQ feeding the final Q has a second consumer (the Add branch), so the
// DQ->Q pair is not a lone sequence and DQQFusion must be rejected. No Convert
// must appear in the compiled graph.
TEST_F(QnnHTPBackendTests, DQQFusion_ExtraConsumer_NoFusion) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CI failed in this testcase.

@@ -0,0 +1,211 @@
// Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

According to the description,

Both fusions previously had only accuracy / EP-assignment coverage; nothing asserted that the fused Convert actually appears in the compiled QNN graph.

Should existing testcases as well move into these new test files?

@@ -0,0 +1,249 @@
// Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR #876 introduces DQQ testcases although it will be reverted in PR #887.
I guess it will eventually re-introduced and please keep track of it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants