A spectre is haunting Unicode

The article explores the history of 'ghost characters' in the Japanese JIS X 0208 encoding standard. These characters were inadvertently created due to cataloging errors and lack of clear source documentation during the digitization of place names.
Why it matters
This highlights the challenges of data integrity and the long-term consequences of errors in foundational digital standards.
In 1978 Japan's Ministry of Economy, Trade and Industry established the encoding that would later be known as JIS X 0208, which still serves as an important reference for all Japanese encodings. However, after the JIS standard was released people noticed something strange - several of the added characters had no obvious sources, and nobody could tell what they meant or how they should be pronounced. Nobody was sure where they came from. These are what came to be known as the ghost characters ( 幽霊文字 ).
For a long time the ghost characters remained an unexplained and mostly forgotten curiosity, but in 1997 an investigation was launched to discover where they had come from. While all characters in the JIS standard were supposed to have a record of their sources, even when it existed it wasn't very specific, typically just listing the document it was sourced from.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in