-
Notifications
You must be signed in to change notification settings - Fork 122
Adding agent files for oneDPL #2668
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
553d8be
d17018a
850a8f1
ada243b
84b555c
ebd226c
ec6f048
e3a9838
223819f
a5ae2cb
37dda02
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
| @../AGENTS.md |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,128 @@ | ||
| # AGENTS.md | ||
|
|
||
| This file provides guidance to AI Agents when working with code in this repository. | ||
|
|
||
|
mmichel11 marked this conversation as resolved.
|
||
| ## Project Overview | ||
|
|
||
| oneDPL (oneAPI Data Parallel Library) is a header-only C++17 library implementing the [oneAPI specification](https://github.com/uxlfoundation/oneAPI-spec/tree/main/source/elements/oneDPL) for parallel algorithms. It provides C++ standard library-like parallel algorithms that work across heterogeneous devices (CPUs, GPUs) using different parallel backends (TBB, DPCPP/SYCL, OpenMP, serial). | ||
|
|
||
| **Key Resources:** | ||
| - [IMPLEMENTATION_DETAILS.md](IMPLEMENTATION_DETAILS.md) - architecture, roadmap and development patterns | ||
| - [cmake/README.md](cmake/README.md) - Complete build system documentation | ||
| - [CONTRIBUTING.md](CONTRIBUTING.md) - Testing requirements and contribution guide | ||
| - [Official Documentation](https://uxlfoundation.github.io/oneDPL) | ||
|
|
||
| ## Quick Start: Building and Testing | ||
|
|
||
| ### Configure Build | ||
| ```bash | ||
| cmake -DCMAKE_CXX_COMPILER=icpx \ | ||
| -DCMAKE_CXX_STANDARD=17 \ | ||
| -DONEDPL_BACKEND=dpcpp \ | ||
| -DCMAKE_BUILD_TYPE=Release \ | ||
| -B build | ||
| ``` | ||
|
|
||
| **Common CMake Variables:** | ||
| - `ONEDPL_BACKEND`: tbb (default), dpcpp, dpcpp_only, omp, serial | ||
| - `CMAKE_BUILD_TYPE`: Debug, Release, RelWithDebInfo, RelWithAsserts | ||
|
|
||
| **Note:** `RelWithAsserts` is Release without `-DNDEBUG`, useful for testing assert-heavy code. | ||
| **Device selection:** For SYCL/DPC++ device selection, use `ONEAPI_DEVICE_SELECTOR` as documented in the [Device Selection](#device-selection) section below. | ||
|
|
||
| ### Build and Run Tests | ||
|
|
||
| ```bash | ||
| # Build specific test | ||
| cmake --build build --target sort.pass | ||
|
|
||
| # Build all tests | ||
| cmake --build build --target build-onedpl-tests | ||
|
|
||
| # Build by category | ||
| cmake --build build --target build-onedpl-algorithm-tests | ||
|
|
||
| # Run specific test | ||
| cd build && ctest -R ^sort.pass$ | ||
|
|
||
| # Run tests by category | ||
| ctest -L ^algorithm$ | ||
|
|
||
| # Run test executable directly | ||
| ./build/test/sort.pass | ||
| ``` | ||
|
danhoeflinger marked this conversation as resolved.
|
||
|
|
||
| **See `cmake/README.md` for complete testing options.** | ||
|
|
||
|
danhoeflinger marked this conversation as resolved.
|
||
| ### Build Documentation | ||
|
|
||
| ```bash | ||
| # Setup documentation | ||
| cd documentation | ||
| python3 -m venv docdpl | ||
| source docdpl/bin/activate | ||
| pip install -r _auxiliary/requirements.txt | ||
|
|
||
| # Generate HTML in build/html | ||
| make html | ||
| ``` | ||
|
|
||
| ## Architecture Quick Reference | ||
|
|
||
| oneDPL uses a three-tier architecture (see `IMPLEMENTATION_DETAILS.md` for details): | ||
|
danhoeflinger marked this conversation as resolved.
|
||
|
|
||
| 1. **Public API Layer** (`include/oneapi/dpl/`) - Standard-like algorithm headers | ||
| 2. **Pattern Layer** (`include/oneapi/dpl/pstl/`, `include/oneapi/dpl/internal/`) - `__pattern_*()` implementations and extensions | ||
| 3. **Backend Layer** (`include/oneapi/dpl/pstl/`) - Parallel execution primitives | ||
|
|
||
| **Key Pattern:** Algorithm → Pattern Function → Backend Primitive | ||
|
|
||
| Example: `std::any_of()` → `__pattern_any_of()` → `__parallel_or()` | ||
|
|
||
| ## Test Organization | ||
|
|
||
| - `test/parallel_api/` - Parallel algorithm tests (algorithm, numeric, memory, ranges) | ||
| - `test/xpu_api/` - Device-side API tests | ||
| - `test/general/` - General functionality | ||
| - `test/kt/` - Kernel templates (experimental) | ||
|
|
||
| Tests use `.pass.cpp` suffix. | ||
|
|
||
| # Code styling | ||
| - Code should be self-documenting. | ||
| - Only use comments when necessary to explain something that is unclear. | ||
| - Do not create short 1-2 line functions only used once unless it improves understandability. | ||
|
|
||
| ## Code Formatting | ||
|
|
||
| **clang-format is required** for all code except tests: | ||
| ```bash | ||
| clang-format -i <file> | ||
| ``` | ||
| Contributors can override clang-format suggestions for readability in exceptional cases. | ||
|
|
||
| When touching existing code, migrate any `::std::` usages to `std::` within the functions touched. | ||
|
|
||
|
danhoeflinger marked this conversation as resolved.
|
||
| ## Device Selection | ||
|
|
||
| SYCL/DPC++ device selection via environment variables: | ||
|
|
||
| ```bash | ||
| # Intel LLVM Compiler >= 2023.1 | ||
| # Uncomment exactly one of the following selectors: | ||
| export ONEAPI_DEVICE_SELECTOR=level_zero:gpu | ||
| # export ONEAPI_DEVICE_SELECTOR=opencl:gpu | ||
| # export ONEAPI_DEVICE_SELECTOR=*:cpu | ||
| ``` | ||
|
|
||
| ## Code Review Guidelines | ||
|
|
||
| - **Formatting-only changes are frowned upon.** PRs should not include changes that are purely cosmetic (whitespace, brace style, etc.) without substantive functional changes. If a file is being modified for functional reasons, incidental formatting fixes in the same area are acceptable, but reformatting unrelated to the PR's purpose should be flagged in reviews and requested to be removed. | ||
|
|
||
|
mmichel11 marked this conversation as resolved.
|
||
| ## Important Notes | ||
|
|
||
| - **Header-only library** - no binary artifacts | ||
| - **C++17 minimum** required; parallel ranges API requires **C++20** | ||
| - **Backend selection is compile-time only** - zero runtime overhead | ||
| - **Two-phase headers** - `*_defs.h` (declarations), `*_impl.h` (implementations) | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Do we want a requirement specifying that all AI-authored commits should be disclosed in the commit message? I think if we list it here, the agent should respect it. It brings up another question about if we should adjust the contribution guide to give user guidance for AI generated code (e.g. all commit should be human reviewed prior to opening a PR, but I think that can be addressed separately).
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Its a good question. I think ultimately the committing user has to be responsible for the code whether or not it was partially or fully generated by an agent. In this case, I'm not so sure that adding something to the commit message is important, but its worth discussing and perhaps making a policy. |
||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
| @AGENTS.md |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,6 @@ | ||
| Use your tool to write directly to files rather than using cat or other workarounds. | ||
|
|
||
| Unless asked to fix or implement something, don't jump to editing code or fixing a problem, | ||
| first explain your plan and check with the user before going forward. | ||
|
|
||
| @AGENTS.md |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,153 @@ | ||
| # oneDPL Implementation Details | ||
|
|
||
| This document describes the high-level architecture and development patterns in oneDPL. | ||
|
|
||
| oneDPL is a header-only library, and has a C++17 minimum requirement. | ||
|
|
||
| ## Three-Tier Design Pattern | ||
|
|
||
| oneDPL uses a layered architecture to provide a single API across multiple backends: | ||
|
|
||
| ### 1. Glue/Public API Layer | ||
| **Location:** `include/oneapi/dpl/` | ||
|
|
||
| Public headers matching C++ standard library naming: | ||
| - `algorithm`, `execution`, `numeric`, `memory`, `ranges`, `iterator` | ||
|
|
||
| Implementation in `glue_*.h` files: | ||
| - Thin wrappers accepting execution policies | ||
| - Dispatch to pattern layer using `__select_backend()` | ||
| - Files split: `*_defs.h` (declarations), `*_impl.h` (implementations) | ||
|
|
||
| ### 2. Pattern Implementation Layer | ||
| **Location:** `include/oneapi/dpl/pstl/` (e.g., `algorithm_impl.h`, `numeric_impl.h`) | ||
|
|
||
| Core algorithm logic in `__pattern_*()` functions: | ||
| - Polymorphic overloads based on: | ||
| - Iterator type (forward, random-access) | ||
| - Execution policy (serial, parallel, hetero) | ||
| - Vectorization capability (`_IsVector` template parameter) | ||
| - Enable single algorithm to serve multiple execution contexts | ||
|
|
||
| Note: `include/oneapi/dpl/internal/` contains some extension implementations (scan-by-segment, binary search, dynamic selection, etc.). | ||
|
|
||
| Example flow: `std::any_of()` → `__internal::__pattern_any_of()` → backend primitives | ||
|
|
||
| ### 3. Backend Implementation Layer | ||
| **Location:** `include/oneapi/dpl/pstl/` | ||
|
|
||
| Backend-specific parallelism primitives: | ||
|
|
||
| **Host Backends:** | ||
| - `parallel_backend_tbb.h` - Intel Threading Building Blocks | ||
| - `parallel_backend_omp.h` - OpenMP pragmas | ||
| - `parallel_backend_serial.h` - Sequential fallback | ||
|
|
||
| **Heterogeneous Backend:** | ||
| - `hetero/dpcpp/parallel_backend_sycl*.h` - SYCL/DPC++ for GPUs/accelerators | ||
|
|
||
| Each backend implements common interface: | ||
| - `__parallel_for()`, `__parallel_reduce()`, `__parallel_scan()`, etc. | ||
| - Compile-time selection via namespace aliasing (`__par_backend`) | ||
|
|
||
| ## Execution Policies | ||
|
|
||
| Policies control algorithm execution strategy: | ||
|
|
||
| | Policy | Execution | Use Case | | ||
| |--------|-----------|----------| | ||
| | `sequenced_policy` (seq) | Sequential | Single-threaded CPU | | ||
|
danhoeflinger marked this conversation as resolved.
|
||
| | `parallel_policy` (par) | Parallel threads | Multi-core CPU with TBB/OpenMP | | ||
| | `parallel_unsequenced_policy` (par_unseq) | Parallel + SIMD | Multi-core CPU with vectorization | | ||
| | `unsequenced_policy` (unseq) | SIMD only | Single-threaded with vectorization | | ||
| | Device policies | GPU/accelerator | SYCL/DPC++ execution | | ||
|
|
||
| Policies map to execution tags: | ||
| - `__serial_tag` - Sequential execution | ||
| - `__parallel_tag<_IsVector>` - Parallel execution with optional vectorization | ||
| - `__hetero_tag<_BackendTag>` - Device execution | ||
|
|
||
| ## Backend Architecture | ||
|
|
||
| ### Backend Selection | ||
|
|
||
| Backends selected at **compile-time only** via CMake `ONEDPL_BACKEND`: | ||
|
|
||
| ```cmake | ||
| -DONEDPL_BACKEND=tbb # TBB (default for most platforms) | ||
| -DONEDPL_BACKEND=dpcpp # SYCL/DPC++ with TBB for host | ||
| -DONEDPL_BACKEND=dpcpp_only # SYCL/DPC++ without TBB | ||
| -DONEDPL_BACKEND=omp # OpenMP | ||
| -DONEDPL_BACKEND=serial # Sequential | ||
| ``` | ||
|
|
||
| Selection logic in `CMakeLists.txt`: | ||
| - When `ONEDPL_BACKEND` is not set, auto-detects SYCL support: defaults to `dpcpp` if SYCL is available, otherwise `tbb` | ||
| - When `ONEDPL_BACKEND` is explicitly set, auto-detection is skipped; the specified backend's dependencies must be available | ||
| - Namespace `__par_backend` points to active backend | ||
|
|
||
| ### SYCL/DPC++ Backend | ||
|
|
||
| **Specialized Kernel Implementations:** | ||
|
|
||
| | File | Purpose | | ||
| |------|---------| | ||
| | `parallel_backend_sycl_for.h` | Kernel launching primitives | | ||
|
danhoeflinger marked this conversation as resolved.
|
||
| | `parallel_backend_sycl_reduce.h` | Reduction operations | | ||
| | `parallel_backend_sycl_*scan*.h` | Scan/prefix-sum operations | | ||
| | `parallel_backend_sycl_merge*.h` | Merge algorithms | | ||
| | `parallel_backend_sycl_radix_sort*.h` | Optimized radix sort | | ||
| | `parallel_backend_sycl_reduce_by_segment.h` | Segmented operations | | ||
|
|
||
| ## Development Patterns | ||
|
|
||
| ### Pattern-Based Implementation | ||
|
|
||
| Standard pattern for implementing algorithms: | ||
|
|
||
| 1. **Public API** (`glue_algorithm_impl.h`) | ||
| - Accepts execution policy + algorithm parameters | ||
| - Dispatches to pattern function | ||
|
|
||
| 2. **Pattern Function** (`algorithm_impl.h`) | ||
| - `__pattern_*()` with overloads for different contexts | ||
| - Selects appropriate brick/backend based on: | ||
| - Iterator category | ||
| - Execution policy | ||
| - Vectorization support | ||
|
|
||
| 3. **Backend Primitive** | ||
| - Actual parallel execution via backend | ||
| - Example: `__par_backend::__parallel_or()` | ||
|
|
||
| ### Brick Functions | ||
|
|
||
| `__brick_*()` functions provide elementary implementations: | ||
| - Used within work units by patterns | ||
| - Enable code reuse between serial/parallel variants | ||
| - Example: `__brick_any_of()` used by both serial and parallel `__pattern_any_of()` | ||
|
|
||
|
|
||
| ## Testing Architecture | ||
|
|
||
| Tests organized by execution model: | ||
|
|
||
| **`test/parallel_api/`** - Tests using host execution policies | ||
| - `algorithm/` - Parallel algorithm tests | ||
| - `numeric/` - Reduction, scan operations | ||
| - `memory/` - Memory algorithms | ||
| - `ranges/` - Range-based API | ||
|
|
||
| **`test/xpu_api/`** - Tests for device-side API | ||
| - C++ standard APIs callable from SYCL kernels | ||
| - Validates device-side iterator support | ||
|
|
||
| **`test/general/`** - General functionality | ||
| - SYCL iterator tests | ||
| - Policy behavior validation | ||
|
|
||
| **`test/kt/`** - Kernel templates (experimental) | ||
| - Hardware-specific optimizations | ||
| - Requires specific device capabilities | ||
|
|
||
| Test naming: `*.pass.cpp` suffix for tests expected to pass. | ||
Uh oh!
There was an error while loading. Please reload this page.