ISO basic Latin alphabet

The ISO basic Latin alphabet is a Latin-script alphabet and consists of two sets of 26 letters, codified in^[1] various national and international standards and used widely in international communication.

The two sets contain the following 26 letters each:^[1]^[2]

ISO basic Latin alphabet
Uppercase Latin alphabet	A	B	C	D	E	F	G	H	I	J	K	L	M	N	O	P	Q	R	S	T	U	V	W	X	Y	Z
Lowercase Latin alphabet	a	b	c	d	e	f	g	h	i	j	k	l	m	n	o	p	q	r	s	t	u	v	w	x	y	z

History

By the 1960s it became apparent to the computer and telecommunications industries in the First World that a non-proprietary method of encoding characters was needed. The International Organization for Standardization (ISO) encapsulated the Latin script in their (ISO/IEC 646) 7-bit character-encoding standard. To achieve widespread acceptance, this encapsulation was based on popular usage. The standard was based on the already published American Standard Code for Information Interchange, better known as ASCII, which included in the character set the 26 × 2 letters of the English alphabet. Later standards issued by the ISO, for example ISO/IEC 8859 (8-bit character encoding) and ISO/IEC 10646 (Unicode Latin), have continued to define the 26 × 2 letters of the English alphabet as the basic Latin script with extensions to handle other letters in other languages.^[1]

Terminology

Name for Unicode block that contains all letters

The Unicode block that contains the alphabet is called "C0 Controls and Basic Latin".

Names for the two subsets

In Unicode 7.0 two subheadings exist:^[3]

"Uppercase Latin alphabet", individual letters contain the string LATIN CAPITAL LETTER in their descriptions
"Lowercase Latin alphabet", individual letters contain the string LATIN SMALL LETTER in their descriptions

Names for the letters

The letters are also contained in "Halfwidth and Fullwidth Forms" FF00 to FFEF^[4]

FF21 Ａ FULLWIDTH LATIN CAPITAL LETTER A
FF41 ａ FULLWIDTH LATIN SMALL LETTER A

Timeline for encoding standards

1865 International Morse Code was standardized at the International Telegraphy Congress in Paris, and was later made the standard by the International Telecommunication Union (ITU)
1950s Radiotelephony Spelling Alphabet by ICAO

Timeline for widely used computer codes supporting the alphabet

1963: ASCII (7-bit character-encoding standard from the American Standards Association, which became ANSI in 1969)
1963/1964: EBCDIC (developed by IBM and supporting the same alphabetic characters as ASCII, but with different code values)
1965-04-30: Ratified by ECMA as ECMA-6^[5] based on work the ECMA's Technical Committee TC1 had carried out since December 1960.^[5]
1972: ISO 646 (ISO 7-bit character-encoding standard, using the same alphabetic code values as ASCII, revised in second edition ISO 646:1983 and third edition ISO/IEC 646:1991 as a joint ISO/IEC standard)
1983: ITU-T Rec. T.51 | ISO/IEC 6937 (a multi-byte extension of ASCII)
1987: ISO/IEC 8859-1:1987 (8-bit character encoding)
- Subsequently, other versions and parts of ISO/IEC 8859 have been published.
Mid-to-late 1980s: Windows-1250, Windows-1252, and other encodings used in Microsoft Windows (some roughly similar to ISO/IEC 8859-1)
1990: Unicode 1.0 (developed by the Unicode Consortium),^[6]^[7] contained in the block "C0 Controls and Basic Latin" using the same alphabetic code values as ASCII and ISO/IEC 646
- Subsequently, other versions of Unicode have been published and it later became a joint ISO/IEC standard as well, as identified below.
1993: ISO/IEC 10646-1:1993, ISO/IEC standard for characters in Unicode 1.1
- Subsequently, other versions of ISO/IEC 10646-1 and one of ISO/IEC 10646-2 have been published. Since 2003, the standards have been published under the name "ISO/IEC 10646" without the separation into two parts.
1997: Windows Glyph List 4

Representation

Hindu-Arabic numerals and letters of the ISO basic Latin alphabet on a 16-segment display.

In ASCII the letters belong to the printable characters and in Unicode since version 1.0 they belong to the block "C0 Controls and Basic Latin". In both cases, as well as in ISO/IEC 646, ISO/IEC 8859 and ISO/IEC 10646 they are occupying the positions in hexadecimal notation 41 to 5A for uppercase and 61 to 7A for lowercase.

Not case sensitive, all letters have code words in the ICAO spelling alphabet and can be represented with Morse code.

Usage

All of the lowercase letters are used in the International Phonetic Alphabet (IPA). In X-SAMPA and SAMPA these letters have the same sound value as in IPA. In Kirshenbaum they have the same value except for the letter r.

Alphabets containing the same set of letters

The list below only includes alphabets that lack:

letters whose diacritical marks make them distinct letters.
multigraphs that constitute distinct letters.

alphabet	diacritic	multigraphs (not constituting distinct letters)	ligatures
Afrikaans alphabet	á, é, è, ê, ë, í, î, ï, ó, ô, ú, û, ý
Catalan alphabet	à, é, è, í, ï, ó, ò, ú, ü, ç
Dutch alphabet	ä, é, è, ë, ï, ö, ü	The digraph ⟨ij⟩ is sometimes considered to be a separate letter. When that is the case, it usually replaces or is intermixed with ⟨y⟩.
English alphabet	-none-	sh, ch, ea, ou, th, ph, ng, zh	æ, œ
French alphabet	à, â, ç, é, è, ê, ë, î, ï, ô, ù, û, ü, ÿ	⟨ai⟩, ⟨au⟩, ⟨ei⟩, ⟨eu⟩, ⟨oi⟩, ⟨ou⟩, ⟨eau⟩, ⟨ch⟩, ⟨ph⟩, ⟨gn⟩, ⟨an⟩, ⟨am⟩, ⟨en⟩, ⟨em⟩, ⟨in⟩, ⟨im⟩, ⟨on⟩, ⟨om⟩, ⟨un⟩, ⟨um⟩, ⟨yn⟩, ⟨ym⟩, ⟨ain⟩, ⟨aim⟩, ⟨ein⟩, ⟨oin⟩, ⟨aî⟩, ⟨eî⟩	æ, œ
German alphabet	ä, ö, ü	⟨sch⟩, ⟨qu⟩, ⟨ch⟩, ⟨ph⟩, ⟨ng⟩, ⟨ie⟩, ⟨ck⟩, ⟨ei⟩, ⟨eu⟩, ⟨äu⟩	ß
Ido alphabet	-none-	⟨qu⟩, ⟨ch⟩, ⟨sh⟩	-none-
Indonesian alphabet	-none-	⟨kh⟩, ⟨ng⟩, ⟨ny⟩, ⟨sy⟩
Interglossa alphabet	-none-
Interlingua alphabet	-none-	⟨qu⟩	-none-
Luxembourgish alphabet	ä, é, ë
Malay alphabet	-none-	⟨gh⟩, ⟨kh⟩, ⟨ng⟩, ⟨ny⟩, ⟨sy⟩	-none-
Occidental alphabet	-none-
Portuguese alphabet	ã, õ, á, é, í, ó, ú, â, ê, ô, à, ç	⟨ch⟩, ⟨lh⟩, ⟨nh⟩, ⟨rr⟩, ⟨ss⟩, ⟨am⟩, ⟨em⟩, ⟨im⟩, ⟨om⟩, ⟨um⟩, ⟨ãe⟩, ⟨ão⟩, ⟨õe⟩	-none-

English is the only major modern European language requiring no diacritics for native words (although a diaeresis is used by some publishers in words such as "coöperation").^[8]^[9]

Note for Portuguese: k, w and y were part of the alphabet until several spelling reforms during the 20th century, the aim of which was to change the etymological Portuguese spelling into an easier phonetic spelling. These letters were replaced by other letters having the same sound: thus psychologia became psicologia, kioske became quiosque, martyr became mártir, etc. Nowadays k, w, and y are only found in foreign words and their derived terms and in scientific abbreviations (e.g. km, byronismo). These letters are considered part of the alphabet again following the 1990 Portuguese Language Orthographic Agreement, which came into effect on January 1, 2009, in Brazil. See Reforms of Portuguese orthography.

Column numbering

The Roman (Latin) alphabet is commonly used for column numbering in a table or chart. This avoids confusion with row numbers using Hindu-Arabic numerals. For example, a 3-by-3 table would contain Columns A, B, and C, set against Rows 1, 2, and 3. If more columns are needed beyond Z (normally the final letter of the alphabet), the column immediately after Z is AA, followed by AB, and so on. This can be seen by scrolling far to the right in a spreadsheet program such as Microsoft Excel or LibreOffice Calc.

These are double-digit "letters" for table columns, in the same way that 10 through 99 are double-digit numbers. The Greek alphabet has a similar extended form that uses such double-digit letters if necessary, but it is used for chapters of a fraternity as opposed to columns of a table.

Such double-digit letters for bullet points are AA, BB, CC, etc., as opposed to the number-like place value system explained above for table columns.

References

1 2 3 "Internationalisation standardization of 7-bit codes, ISO 646". Trans-European Research and Education Networking Association (TERENA). Retrieved 2010-10-03.
↑ "RFC1815 – Character Sets ISO-10646 and ISO-10646-J-1". Retrieved 2010-10-03.
↑ "CO Controls and Basic Latin" (PDF). Unicode.org. Retrieved 2016-08-08.
↑ "Halfwidth and Fullwidth Forms" (PDF). Unicode.org. Retrieved 2016-08-08.
1 2 Standard ECMA-6: 7-Bit Coded Character Set (PDF) (5th ed.). Geneva, Switzerland: European Computer Manufacturers Association (Ecma). March 1985. Archived (PDF) from the original on May 29, 2016. Retrieved 2016-05-29. The Technical Committee TC1 of ECMA met for the first time in December 1960 to prepare standard codes for Input/Output purposes. On April 30, 1965, Standard ECMA-6 was adopted by the General Assembly of ECMA.
↑ "Unicode character database". The Unicode Standard. Retrieved 2013-03-22.
↑ The Unicode Standard Version 1.0, Volume 1. Addison-Wesley Publishing Company, Inc. 1990. ISBN 0-201-56788-1.
↑ As an example, an article containing a diaeresis in "coöperate" and a cedilla in "façades" as well as a circumflex in the word "crêpe" (Grafton, Anthony (2006-10-23). "Books: The Nutty Professors, The history of academic charisma". The New Yorker. )
↑ "The New Yorker's odd mark — the diaeresis"

Latin script

Alphabets (list)

Classical Latin alphabet
ISO basic Latin alphabet
phonetic alphabets
- International Phonetic Alphabet
- X-SAMPA
spelling alphabets

Letters (list)

Letters of the ISO basic Latin alphabet
Aa	Bb	Cc	Dd	Ee	Ff	Gg	Hh	Ii	Jj	Kk	Ll	Mm	Nn	Oo	Pp	Qq	Rr	Ss	Tt	Uu	Vv	Ww	Xx	Yy	Zz

Multigraphs

Digraphs	ch cz dž dz gh ij ll ly nh ny sh sz th
Trigraphs	dzs eau
Tetragraphs	ough
Pentagraphs	tzsch

Keyboard layouts (list)

Standards

Lists

Character encodings
Early telecommunications	ASCII ISO/IEC 646 ISO/IEC 6937 T.61 BCDIC Baudot code Morse code Telegraph code Wabun code Special telegraphy codes Non-Latin Chinese Cyrillic Needle telegraph codes
ISO/IEC 8859	-1 -2 -3 -4 -5 -6 -7 -8 -9 -10 -11 -12 -13 -14 -15 -16
Bibliographic use	ANSEL ISO 5426 / 5426-2 / 5427 / 5428 / 6438 / 6861 / 6862 / 10585 / 10586 / 10754 / 11822 MARC-8
National standards	ArmSCII BraSCII CNS 11643 ELOT 927 GOST 10859 GB 18030 HKSCS ISCII JIS X 0201 JIS X 0208 JIS X 0212 JIS X 0213 KOI-7 KPS 9566 KS X 1001 PASCII SI 960 TIS-620 TSCII VISCII YUSCII
EUC	CN JP KR TW
ISO/IEC 2022	CN JP KR CCCII
MacOS code pages ("scripts")	Arabic Celtic CentEuro ChineseSimp / EUC-CN ChineseTrad / Big5 Croatian Cyrillic Devanagari Dingbats Esperanto Farsi Gaelic Greek Gujarati Gurmukhi Hebrew Iceland Japanese / ShiftJIS Korean / EUC-KR Latin-1 Roman Romanian Sámi Symbol Thai / TIS-620 Turkish Ukrainian
DOS code pages	100 111 112 113 151 152 161 162 163 164 165 166 210 220 301 437 449 489 620 667 668 707 708 709 710 711 714 715 720 721 737 768 770 771 772 773 774 775 776 777 778 790 850 851 852 853 854 855/872 856 857 858 859 860 861 862 863 864/17248 865 866/808 867 868 869 874/1161/1162 876 877 878 881 882 883 884 885 891 895 896 897 898 899 900 903 904 906 907 909 910 911 926 927 928 929 932 934 936 938 941 942 943 944 946 947 948 949 950/1370 951 966 991 1034 1039 1040 1041 1042 1043 1044 1046 1086 1088 1092 1093 1098 1108 1109 1114 1115 1116 1117 1118 1119 1125/848 1126 1127 1131/849 1139 1167 1168 1300 1351 1361 1362 1363 1372 1373 1374 1375 1380 1381 1385 1386 1391 1392 1393 1394 Kamenický Mazovia CWI-2 KOI8 MIK Iran System
IBM AIX code pages	367 371 806 813 819 895 896 912 913 914 915 916 919 920 921/901 922/902 923 952 953 954 955 956 957 958 959 960 961 963 964 965 970 971 1004 1006 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1029 1036 1089 1111 1124 1129/1163 1133 1350 1382 1383
IBM Apple MacIntosh emulations	1275 1280 1281 1282 1283 1284 1285 1286
IBM Adobe emulations	1038 1276 1277
IBM DEC emulations	1020 1021 1023 1090 1100 1101 1102 1103 1104 1105 1106 1107 1287 1288
IBM HP emulations	1050 1051 1052 1053 1054 1055 1056 1057 1058
Windows code pages	CER-GS 874/1162 (TIS-620) 932/943 (Shift JIS) 936/1386 (GBK) 950/1370 (Big5) 949/1363 (EUC-KR) 1169 1174 Extended Latin-8 1200 (UTF-16LE) 1201 (UTF-16BE) 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1261 1270 54936 (GB18030)
EBCDIC code pages	1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37/1140 38 39 40 251 252 254 256 257 258 259 260 264 273/1141 274 275 276 277/1142 278/1143 279 280/1144 281 282 283 284/1145 285/1146 286 287 288 289 290 293 297/1147 298 300 310 320 321 322 330 351 352 353 355 357 358 359 360 361 363 382 383 384 385 386 387 388 389 390 391 392 393 394 395 410 420/16804 421 423 424/8616/12712 425 435 500/1148 803 829 833 834 835 836 837 838/838 839 870/1110/1153 871/1149 875/4971/9067 880 881 882 883 884 885 886 887 888 889 890 892 893 905 918 924 930/1390 931 933/1364 935/1388 937/1371 939/1399 1001 1002 1003 1005 1007 1024 1025/1154 1026/1155 1027 1028 1030 1031 1032 1033 1037 1047 1068 1069 1070 1071 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1087 1091 1097 1112/1156 1113 1122/1157 1123/1158 1130/1164 1132 1136 1137 1150 1151 1152 1159 1165 1166 1278 1279 1303 1364 1376 1377 JEF KEIS
Platform specific	Acorn Adobe Standard Apple II ATASCII Atari ST BICS Casio calculators CDC CPC DEC Radix-50 DEC MCS/NRCS DG International ELWRO-Junior FIELDATA GEM GEOS GSM 03.38 HP Roman Extension HP Roman-8 HP Roman-9 HP calculators LICS LMBCS MSX NEC APC NeXT PCW PETSCII Sharp calculators TI calculators TRS-80 Ventura International Ventura Symbol WISCII XCCS ZX80 ZX81 ZX Spectrum
Unicode / ISO/IEC 10646	UTF-1 UTF-7 UTF-8 UTF-16 (UTF-16LE/UTF-16BE) / UCS-2 UTF-32 (UTF-32LE/UTF-32BE) / UCS-4 UTF-EBCDIC GB 18030 BOCU-1 CESU-8 SCSU
Miscellaneous code pages	ABICOMP APL ARIB STD-B24 Cork HZ INIS INIS-8 Johab LY1 OML OMS OT1 SEASCII TACE16 TRON UTF-5 UTF-6 WTF-8
Related topics	Code page Control character (C0 C1) CCSID Character encodings in HTML Charset detection Han unification Hardware ISO 6429/IEC 6429/ANSI X3.64 Mojibake
Character sets

This article is issued from Wikipedia. The text is licensed under Creative Commons - Attribution - Sharealike. Additional terms may apply for the media files.