Module talk:Unicode data

About RTL

I am researching RTL scripts. I met this:

A

0xa9 -- LATIN CAPITAL LETTER A

Latn

is_rtl: false

ث

0x062B -- ARABIC LETTER THEH [1]

Arab

is_rtl: false

ש

0x05E9 -- HEBREW LETTER SHIN [2]

Hebr

is_rtl: false

ߖ

0x07D6 -- NKO LETTER JA [3]

Nkoo

is_rtl: false

I'd expect the Arab, Hebr, Nkoo characters to be rtl=true. Am I misunderstanding something? @Erutuon: -DePiep (talk) 20:58, 9 January 2021 (UTC)[reply]

@DePiep: The invocation {{#invoke:Unicode data|is|rtl|05E9}} checks whether the literal characters 05E9 are right-to-left. To check the right-to-leftness of the Hebrew character, put in the literal character or a HTML character reference: {{#invoke:Unicode data|is|rtl|ש}} or {{#invoke:Unicode data|is|rtl|ש}}. #invoke:Unicode data|is|rtl as well as #invoke:Unicode data|is|valid_pagename and #invoke:Unicode data|is|Latin interpret their arguments as strings rather than code points in hexadecimal because the corresponding functions in the module take strings. (They could take hexadecimal arguments if someone edited the module to add another parameter to tell them to interpret their argument this way.) — Eru·tuon 01:02, 10 January 2021 (UTC)[reply]

@Erutuon: Thanks, will work for me. Great module! (Second code example is {{#invoke:Unicode data|is|rtl|ש}}). -DePiep (talk) 17:28, 10 January 2021 (UTC)[reply]

The four characters, is_rtl:

using &#x...; false

using &#x...; true

-DePiep (talk) 20:23, 10 January 2021 (UTC)[reply]

is_pagename

Resolved

In the function is_pagename, does "pagename" stand for "blockname"? Or wider? -DePiep (talk) 05:17, 27 March 2022 (UTC)[reply]

Resolved: refers to "valid WP pagename", related to WP:NCTR invalid title characters like "#". -DePiep (talk) 11:34, 27 March 2022 (UTC)[reply]

Missing documentation: Hangul, Aliases

I am developing the documentation, especially in Module:Unicode data § List of functions. To completify, can someone point out how or where the data /aliases and /Hangul can be retrieved (implementation)? DePiep (talk) 11:39, 27 March 2022 (UTC)[reply]

is_RTL check?

About U+0634 ش ARABIC LETTER SHEEN [4]:

{{#invoke:Unicode data |is|rtl|0x0634}} → false

I expect true (is_rtl), right? -DePiep (talk) 23:00, 28 March 2022 (UTC)[reply]

Solved: enter the character <ش >, not the U+hex:

{{#invoke:Unicode data |is|rtl|ش }} → true

DePiep (talk) 05:26, 1 June 2022 (UTC)[reply]

Edit request 20 November 2023

This edit request has been answered. Set the |answered= parameter to no to reactivate your request.

Description of suggested change: the module code says "-- No image data modules on Wikipedia yet."

We have them now. Can this be enabled? — Alexis Jazz (talk or ping me) 05:37, 20 November 2023 (UTC)[reply]

Can you sandbox the code? — Martin (MSGJ · talk) 12:46, 20 November 2023 (UTC)[reply]

MSGJ, I don't speak Lua.. I edited Module:Unicode data/sandbox to sync with the current version and I uncommented the block.
{{#invoke:Unicode data/sandbox|lookup|image|0xA9}} returns Unicode 0x00A9.svg (File:Unicode 0x00A9.svg) so I think this works? — Alexis Jazz (talk or ping me) 21:19, 20 November 2023 (UTC)[reply]

Done I'm not sure I agree with your importing of so many modules from other wikis, but in any event there was never any good reason to comment out that code as opposed to just letting uses of it fail. * Pppery * _{it has begun...} 21:36, 22 November 2023 (UTC)[reply]

Edit request 20 April 2024

This edit request has been answered. Set the |answered= parameter to no to reactivate your request.

Description of suggested change: Creation of p.is_noncharacter() as a separate function

Diff:

Eievie (talk) 20:48, 20 April 2024 (UTC)[reply]

Done * Pppery * _{it has begun...} 15:22, 21 April 2024 (UTC)[reply]

Edit request 1 January 2025

This edit request has been answered. Set the |answered= parameter to no to reactivate your request.

Description of suggested change:

Allow looking up the kCantonese Unihan property. As an example, {{#invoke:Unicode data/sandbox|lookup|kCantonese|20EB6}} returns "naap6".

Diff:

function p.lookup_kCantonese(codepoint)
	local data = loader[('Unihan/kCantonese/%02X'):format(floor(codepoint / 0x1000))]
	if data then
		return data[codepoint]
	end
end

Northern Moonlight 03:54, 1 January 2025 (UTC)[reply]

Done * Pppery * _{it has begun...} 23:05, 13 January 2025 (UTC)[reply]

Edit request 15 June 2025

This edit request has been answered. Set the |answered= parameter to no to reactivate your request.

Description of suggested change: Reorder the name_hooks table so its entries are sorted in codepoint order. binary_range_search assumes the entries are sorted in this way currently and therefore does not work correctly. {{unichar}} is currently broken by this bug as can be seen in CJK Unified Ideographs Extension I § Background. Specifically U+2ED9D 𮶝 CJK UNIFIED IDEOGRAPH-2ED9D and U+2EDE0 𮷠 CJK UNIFIED IDEOGRAPH-2EDE0 incorrectly appear as reserved. I have made the change in the sandbox.

Diff: ~~See comparison of sandbox with main~~ Warudo (talk) 12:20, 15 June 2025 (UTC)[reply]

--Warudo (talk) 13:56, 15 June 2025 (UTC)[reply]

Done in Special:Diff/1296263621, thank you. U+2ED9D and U+2EDE0 are now shown correctly. I've also added a test at Template:Unichar/testcases#U+2ED9D – grass radical to show the effect. —⁠andrybak (talk) 22:57, 18 June 2025 (UTC)[reply]

Edit request 29 July 2025

This edit request has been answered. Set the |answered= parameter to no to reactivate your request.

Description of suggested change: Add Variation Selectors (not to be confused with Variation Selectors Supplement) to the name_hooks list. This fixes Template_talk:Unichar#c-Great_Brightstar-20250729153400-Some_character_names_are_not_found_by_the_template properly. (I've added the characters to c:Data:Unicode_data/names/00F.tab but that is a hack and should be reverted once the proper fix is done here.) I've provided the code in the Sandbox which was copied over from wikt:Module:Unicode data which handles this correctly.

Diff:

Warudo (talk) 16:17, 29 July 2025 (UTC)[reply]

Done * Pppery * _{it has begun...} 16:16, 2 August 2025 (UTC)[reply]

@Pppery: I'm sorry for opening this again but I made a copy paste error from Wiktionary which means that the fix failed. The new lines must be added after the CJK compatibility ideographs instead of before them so please make this change:

I missed this because the temporary fix in Commons masked the problem. I can confirm that this time I tested the change properly by reverting my fix in commons first. Warudo (talk) 16:39, 2 August 2025 (UTC)[reply]

Done * Pppery * _{it has begun...} 16:52, 2 August 2025 (UTC)[reply]