Help:Multilingual support
From Wikipedia, the free encyclopedia
Articles on the English Wikipedia may contain words or texts written in different languages and scripts. To be able to correctly view and edit there articles requires that you have the appropriate fonts installed and to have correctly configured your operating system and browser. This guide will help you to do so.
Contents |
[edit] Overview
[edit] Unicode
Articles on Wikipedia are encoded using Unicode (specifically UTF-8)[1], an industry standard designed to allow text and symbols from all of the writing systems of the world to be consistently represented and manipulated by computers. Because Unicode is backwards compatible with ASCII and Latin-1, and most modern browsers have at least basic Unicode support, most users will will experience little difficulty reading and editing Wikipedia.
For older browsers, MediaWiki, the Wikipedia software, serves the wikitext in a safe mode upon editing. Characters that cannot be represented in ASCII are temporarily converted to hexadecimal character references, looking like ሴ. Existing hexadecimal character references get an additional leading zero so they are not converted to actual characters when the page is saved, and look like ሴ. Likewise, to create a hexadecimal character reference in safe mode, not the character itself, a leading zero should be added. One can check whether safe mode is used by editing this section. If M looks like M rather than M, safe mode is used.
[edit] Scripts
[edit] East Asian
[edit] Indic
The following table compares how a correctly enabled computer would render the following scripts with how your computer renders them:
[edit] Special cases
[edit] Esperanto
in edit box | in database and output |
---|---|
S | S |
Sx | Ŝ |
Sxx | Sx |
Sxxx | Ŝx |
Sxxxx | Sxx |
Sxxxxx | Ŝxx |
Mediawiki installations configured for Esperanto use UTF-8 for storage and display. However when editing the text is converted to a form that is designed to be easier to edit with a standard keyboard.
The characters for which this applies are: Ĉ, Ĝ, Ĥ, Ĵ, Ŝ, Ŭ, ĉ, ĝ, ĥ, ĵ, ŝ, ŭ. you may enter these directly in the edit box if you have the facilities to do so. However when you edit the page again you will see them encoded as Sx. This form is referred to as "x-sistemo" or "x-kodo". In order to preserve round trip capability when one or more x's follow these characters or their non-accented forms (A, G, H, J, S, U, c, g, h, j, s, u), the number of x's in the edit box is double the number in the actual stored article text.
For example, the interlanguage link [[en:Luxury car]] to en:Luxury car has to be entered in the edit box as [[en:Luxxury car]] on eo:. This has caused problems with interwiki update bots in the past.
[edit] Romanian
The Romanian alphabet contain an S-comma (Ș ș) and T-comma (Ț ț). These character were added to Unicode 3.0 at the request of the Romanian stadadization institute. Font support for these characters is poor, so the Romanain Wikipedia represents these letters with an S-cedilla (Ş ş) and T-cedilla (Ţ ţ) instead.[2]
[edit] Notes
- ^ Until June 2005, when MediaWiki 1.5 came into use on the Wikimedia projects, articles on the English Wikipedia where encoded using ISO/IEC 8859-1 (although the additional characters from the Windows-1252 character set where used in practised.) All characters from the ISO/IEC 10646 Universal Character Set could be accessed through numerical entities, as specified by the HTML 4.01 specification. Since, nearly all pages have been converted to use Unicode directly.
- ^ See also ro:Wikipedia:Diacritice