Skip to content

fix(instructor): strip invisible Unicode chars from batch-enroll lists - #101

Merged
AliAlfaifi merged 1 commit into
open-release/teak.nelpfrom
fix/batch-enroll-strip-invisible-chars
Oct 6, 2026
Merged

AliAlfaifi merged 1 commit into
open-release/teak.nelpfrom
fix/batch-enroll-strip-invisible-chars

Conversation

@AliAlfaifi

Copy link
Copy Markdown

Bug (seen in prod)

Batch enrollment showed "The following email addresses and/or usernames are invalid:" with an empty bullet.

_split_input_list splits on [\n\r\s,], and Python's \s does not match Unicode format characters (category Cf: U+200F RLM, U+200E LRM, U+200B ZWSP, U+200C/D, U+2060, U+FEFF BOM, U+061C ALM, U+202A–E, U+2066–9). Lists copied from Arabic Excel/Word/WhatsApp carry them, so:

  • a token made only of them survives the split → "invalid" with nothing visible;
  • a real email with one attached (user@example.com + RLM) is reported invalid although it looks correct, and that learner is silently not enrolled.

Fix

Remove every Cf character before splitting, in both copies of the helper (instructor/views/api.py and instructor/views/serializer.py). Format characters never belong in an email, username or National ID, so this is correct for every caller: batch enrollment, CCX coach enrollment, beta testers, access modification. # NELC: marked.

Tests

17 new ddt cases in TestInstructorAPIHelpers (each invisible char alone vanishes; email/username/ID with attached marks are cleaned; Arabic-mixed input); they fail without the fix (17 failed / 4 passed). Container run (overhangio/openedx:20.0.5): instructor test_api.py, test_enrollment.py, ccx → 506 passed, 9 skipped (existing CCX perf tests). Repo pylint (3.3.6 + edx-lint): clean on all touched files.

Stacked under the National ID PR, which targets this branch.

🤖 Generated with Claude Code

…ll identifier lists

_split_input_list splits on [\n\r\s,], and Python's \s does not match Unicode
format characters (category Cf: RLM/LRM, ZWSP, ZWNJ/ZWJ, WORD JOINER, BOM, ALM,
bidi embeddings/isolates). Lists pasted from Arabic Excel/Word carry them, so a
token made only of such characters survived as a blank "invalid identifier" and
a real email with an attached RLM was reported invalid.

Remove all Cf characters before splitting, in both copies of the helper
(views/api.py: batch enrollment + CCX; views/serializer.py: beta testers /
access modification). They never belong in an email, username or national ID.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@AliAlfaifi
AliAlfaifi merged commit 4a3b409 into open-release/teak.nelp Oct 6, 2026
47 checks passed
@AliAlfaifi
AliAlfaifi deployed to open-release/teak.nelp October 6, 2026 20:51 — with GitHub Actions Active
@AliAlfaifi
AliAlfaifi deployed to open-release/teak.nelp October 6, 2026 20:52 — with GitHub Actions Active

This branch was successfully deployed

1 active deployment
open-release/teak.nelp — 9d9ca99d Deployed Oct 6, 2026 by AliAlfaifi via create-jira-issue / create_jira_issue #39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant