Changelog: 0.1.x¶
Archived verbatim from CHANGELOG.md. Newer releases are in the changelog; the archive index lists every older series.
[0.1.5] — 2026-03-27¶
Added¶
- Reverse transliteration:
transliterate(text, target="ru")converts Latin → native script for Russian, Ukrainian, and Greek. PHF tables generated at build time from inverted language TSV data. - Toned pinyin:
transliterate("北京", tones=True)returns"běi jīng"with tone marks. Toned readings sourced from UnihankMandarinfield for all 20,924 CJK Unified Ideographs. - ISO 9:1995 scholarly Cyrillic:
transliterate(text, strict_iso9=True)for scholarly romanization. GOST R 7.0.34 variant viagost7034=True. - Japanese Kunrei-shiki (
lang="ja-kunrei"): alternative romanization profile, bringing total language count to 65. - Ancient scripts: Coptic, Gothic, Old Italic, Runic, Ogham transliteration tables.
- CLI short aliases:
t(transliterate),s(slugify),n(normalize),p(pipeline),d(demojize) — e.g.translit t "café". - CLI
--targetflag:translit t --target ru "Moskva"for reverse transliteration. - CLI
--tones,--strict-iso9,--gost7034flags for transliterate subcommand. - CLI
--langflag for slugify subcommand. console_scriptsentry point:translitcommand available afterpip install translit-rs.docs/cli.md: comprehensive CLI documentation with piping, exit codes, examples.- Links section in README.md and docs/index.md for RTD ↔ GitHub cross-references.
Changed¶
transliterate()API unified:reverse_transliterate()merged intotransliterate()viatargetparameter. Old function removed.transliterate_implRust signature now takes 7 arguments (addedtones: bool).- Updated benchmark numbers after
tonesparameter addition (15–46% regression in transliteration hot path due to additional branch; throughput now 450M chars/sec Latin, 130M chars/sec Cyrillic). - Performance documentation updated across 4 files to reflect current benchmark results.
Fixed¶
- clippy
format_push_stringlint inbuild.rs— replacedpush_str(&format!())withwrite!(). - clippy
unreadable_literalin PHF-generatedreverse_translit_phf.rs— suppressed via inner attribute insrc/reverse.rs. - All 219 integration test call sites updated for 7-argument
transliterate_impl.
[0.1.4] — 2026-03-25¶
Added¶
lang="auto"script-based language detection: Whenlang="auto"is passed totransliterate(),slugify(),TextPipeline,Slugifier, or any other call site, the library detects the dominant non-Latin script in the input and maps it to a default language code automatically. Maps 28 scripts to language codes (e.g. Cyrillic→ru, Han→zh, Hiragana/Katakana→ja, Thai→th). Zero overhead forlang=Noneor explicit lang codes.LANG_AUTOconstant ("auto") intranslit._enums.- Georgian transliteration (
lang="ka"): 114 TSV entries covering Mkhedruli, Mtavruli, and supplement ranges. BGN/PCGN national romanization. - Armenian transliteration (
lang="hy"): 86 TSV entries covering uppercase, lowercase, and 5 ligatures (U+FB13–FB17). BGN/PCGN romanization. - Sinhala transliteration (
lang="si"): 90 TSV entries. Extended Indic Brahmic engine range from0x0900..=0x0D7Fto0x0900..=0x0DFFwith dedicatedsinhala_char_role()function for Sinhala-specific offsets. - Thai transliteration (
lang="th"): 87 TSV entries using RTGS romanization. NewScriptClass::Taiwith tone-mark stripping and cancellation handling. - Lao transliteration (
lang="lo"): 67 TSV entries using BGN/PCGN romanization. Shares Tai engine with Thai via offset masking. - Ethiopic transliteration (
lang="am"): 307 TSV entries for Ge'ez alphasyllabary (34 consonant bases × 7 vowel orders + labialized forms + digits). Pure data addition — no engine changes needed. - Myanmar transliteration (
lang="my"): 89 TSV entries. Newmyanmar_char_role()for Brahmic engine with virama (U+1039) and asat (U+103A) support. Medials (U+103B–103E) classified as dependent vowels. - Khmer transliteration (
lang="km"): 110 TSV entries. Newkhmer_char_role()for Brahmic engine with coeng (U+17D2) as virama. All consonants normalized to inherent 'a' regardless of series. - Tibetan transliteration (
lang="bo"): 147 TSV entries. Newtibetan_char_role()for Brahmic engine with halanta (U+0F84) and subjoined consonants (U+0F90–0FBC). - Unicode range constants:
TIBETAN(0x0F00–0x0FFF),MYANMAR(0x1000–0x109F),KHMER(0x1780–0x17FF) insrc/unicode_ranges.rs. - Comprehensive test coverage: example-based tests for all 9 new scripts, property-based tests (hypothesis + proptest), multi-script mixture tests.
- Built-in language count: 51 → 60.
Changed¶
is_indic()extended to include Tibetan, Myanmar, and Khmer ranges for Brahmic abugida processing.indic_char_role()dispatches to script-specific functions for Sinhala, Tibetan, Myanmar, and Khmer codepoint ranges.
[0.1.3] — 2026-03-25¶
Added¶
strip_controlandstrip_zero_widthnow work as independent pipeline steps without requiringcollapse_whitespace=True. Previously they were silently ignored whencollapse_whitespacewas disabled.strip_control_chars()andstrip_zero_width_chars()standalone Rust functions for filtering without whitespace collapsing.decimalandhexadecimalflags inSlugConfigare now functional. Settingdecimal=Falsepreserves&#NNN;entities;hexadecimal=Falsepreserves&#xHHH;entities. Previously these flags were accepted but silently ignored.- Rust integration tests:
tests/integration_emoji.rs(10 tests),tests/integration_slugify.rs(20 tests),tests/integration_transliterate.rs(21 tests),tests/integration_whitespace.rs(12 tests).
Changed¶
TextPipelineparametersstrip_controlandstrip_zero_widthchanged frombool(defaultTrue) tobool | None(defaultNone). WhenNone, they inherit fromcollapse_whitespace—Trueifcollapse_whitespace=True,Falseotherwise. Set explicitly toTruefor standalone use withoutcollapse_whitespace. This is backward compatible: existing code that passescollapse_whitespace=Truegets the same behavior as before.steps()now reportsstrip_controlandstrip_zero_widthas separate entries when active, giving full visibility into pipeline behavior.- Pipeline step order updated:
normalize → confusables → demojize → strip_accents → transliterate → fold_case → strip_control → strip_zero_width → collapse_whitespace. - Migrated from
once_celltostd::sync::LazyLock/OnceLock; MSRV bumped to 1.80. Removedonce_celldependency. needs_cjk_space()match arm tightened from wildcard_to explicitIdeograph | Hangul | Kanato match the call-siteis_cjkguard.
Fixed¶
decode_entities()corrupting multi-byte UTF-8 characters (BUG-1). The function usedbytes[i] as charwhich treated each continuation byte as a separate Latin-1 codepoint (e.g.café→café). Now advances by full UTF-8 characters.decode_numeric_entity_skip()panicking on malformed&#followed by multi-byte UTF-8 (BUG-2). The skip function walked through continuation bytes looking for;, landing inside a multi-byte character. Now stops at the first non-ASCII byte.
Performance¶
- ASCII fast-path in
demojize_implanddemojize_rust: pure-ASCII text returns immediately withoutVec<char>allocation or emoji scanning. filter_stopwordsreplaced intermediateVec<_>+.join()with a pre-allocatedStringfold, removing one allocation per slugify call.
[0.1.2] — 2026-03-25¶
Added¶
- Python 3.14 support (classifier and CI test matrix).
ruff check --fixpre-commit hook for automatic lint fixing.- CI publish workflow using
pypa/gh-action-pypi-publishwith OIDC trusted publishers. - Multi-platform wheel builds: Linux (x86_64, aarch64), macOS (Intel, ARM64), Windows.
steps()method on_TextPipelinetype stub.
Changed¶
- Resolved all clippy pedantic warnings instead of suppressing them — reduced
lint suppressions from 48 to 22 (remaining are genuine PyO3 constraints).
Fixes include: combined identical match arms, replaced manual counters with
.enumerate(), moved item declarations before statements, usedclone_into(), merged identical branches, fixed doc comment formatting. - Widened
stopwordsandreplacementstype stubs from stricttuple/listtoSequencefor better mypy compatibility. - Applied
ruff formatto all Python source and test files. - Switched docs publish from deprecated
maturin uploadtopypa/gh-action-pypi-publish. - macOS Intel wheels now cross-compiled on ARM64 runner (macos-14) instead of deprecated macos-13.
- CI doctests now run against installed package (not source tree) with explicit
shell: bashfor Windows compatibility.
Fixed¶
TextPipeline.explain()doctest: output format isnormalize (NFC)notnormalize (form=NFC).from __future__ import annotationsplacement in test files (must follow module docstring, not precede it).- Malformed HTML entity test expectation:
decode_entities("&#xyz;")correctly returns"", not"yz;". - Rust benchmark CI: target
bench_corebinary explicitly to avoid passing Criterion flags to the test harness. - Ruff lint fixes: unsorted imports in
test_encoding.py, unused importis_mixed_scriptintest_security_invariants.py. - Read the Docs trigger workflow: simplified curl status handling, graceful
warning when
RTD_TOKENis missing. - Removed incorrect PyPy classifier (abi3 is CPython-only).
[0.1.1] — 2026-03-25¶
Added¶
src/unicode_ranges.rs— named constants for all Unicode codepoint ranges used by the library, eliminating magic numbers scattered across modules.tests/test_concurrency.py— concurrent access tests forLANG_TABLESandHANGUL_CACHE, plus malformed Unicode input tests.- Code coverage reporting in CI (
pytest-cov, XML report uploaded as artifact). CLOCK$,KEYBD$,SCREEN$,COM0,LPT0added to Windows reserved filename list.casefold()alias forfold_case()— matchesstr.casefold()naming.remove_accents()alias forstrip_accents()— matches sklearn/ML ecosystem naming.- Compatibility parameter aliases:
replacement_text/max_lenonsanitize_filename()(pathvalidate),greedy/preferred_aliasesonis_confusable()(confusable_homoglyphs),delimitersondemojize()(emoji library). - Complete API documentation for 19 previously undocumented exported functions:
precompiled pipelines, grapheme clusters, encoding detection,
Textbuilder,is_safe_hostname,demojize,strip_bidi,EmojiProviderprotocol. - Three new API reference pages: Precompiled Pipelines, Grapheme Clusters, Encoding.
- "Guides by role" section in
docs/index.mdandREADME.md. - Performance section in
README.mdwith benchmark numbers. Scriptenum documentation expanded from 28 to all 41 members.
Changed¶
transliterate_implrefactored: capacity estimation extracted toestimate_capacity(), character classification toclassify_char(), and CJK spacing logic toneeds_cjk_space().- All
RwLockaccesses now recover from lock poisoning using.unwrap_or_else(|e| e.into_inner())instead of silently falling through. - Lambda closures in
_compat.pyreplaced with named inner functions for clarity. emoji.rswrite!()call no longer uses.unwrap()(infallible, documented with a// SAFETYcomment).- MkDocs theme switched from
materialtoreadthedocs. - All documentation references updated from "unirust" to "translit".
- Development status promoted from Alpha to Beta.
- Package renamed from
translittotranslit-rson PyPI (interim until PEP 541 grants thetranslitname). Python import remainsimport translit.
Fixed¶
- Type stub
_text.pyiimported from wrong module name (unirust→translit). - Type stub
_translit.pyimissingmin_confidenceparameter on_decode_to_utf8. - Type stub
_text.pyimissinggrapheme_split,grapheme_truncate,catalog_keymethods. security_clean()pipeline step order corrected in 5+ locations: strip_bidi runs before collapse_whitespace (matching Rust implementation).catalog_key()step order corrected: transliterate before strip_accents.- Stale PyO3 boundary overhead corrected from ~4µs to ~240ns in docs and code comments.
Deprecated¶
translit._compatawesome-slugify compatibility layer (Slugify,UniqueSlugify,slugify_*instances) — planned removal in v1.0.
[0.1.0] — 2026-01-01¶
Added¶
- Initial release.
- Unicode transliteration for 60 language profiles.
- Slugification, normalization, confusable detection, filename sanitization.
- Emoji demojization with ZWJ sequence support.
- Backward-compatible layers for Unidecode and awesome-slugify.