Probing the Statistical Properties of Unknown Texts: Application to the Voynich Manuscript

Amancio, D.R.; Altmann, E.G.; Rybski, D.; Oliveira Jr., O.N.; da Costa, L.F.

doi:https://doi.org/10.34657/3901

Probing the Statistical Properties of Unknown Texts: Application to the Voynich Manuscript

dc.bibliographicCitation.firstPage	e67310	eng
dc.bibliographicCitation.issue	7	eng
dc.bibliographicCitation.journalTitle	PLoS ONE	eng
dc.bibliographicCitation.volume	8	eng
dc.contributor.author	Amancio, D.R.
dc.contributor.author	Altmann, E.G.
dc.contributor.author	Rybski, D.
dc.contributor.author	Oliveira Jr., O.N.
dc.contributor.author	da Costa, L.F.
dc.date.accessioned	2020-08-01T15:36:09Z
dc.date.available	2020-08-01T15:36:09Z
dc.date.issued	2013
dc.description.abstract	While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed on the interdependence between syntactic and semantic factors. In this study we propose a framework for determining whether a text (e.g., written in an unknown alphabet) is compatible with a natural language and to which language it could belong. The approach is based on three types of statistical measurements, i.e. obtained from first-order statistics of word properties in a text, from the topology of complex networks representing texts, and from intermittency concepts where text is treated as a time series. Comparative experiments were performed with the New Testament in 15 different languages and with distinct books in English and Portuguese in order to quantify the dependency of the different measurements on the language and on the story being told in the book. The metrics found to be informative in distinguishing real texts from their shuffled versions include assortativity, degree and selectivity of words. As an illustration, we analyze an undeciphered medieval manuscript known as the Voynich Manuscript. We show that it is mostly compatible with natural languages and incompatible with random texts. We also obtain candidates for keywords of the Voynich Manuscript which could be helpful in the effort of deciphering it. Because we were able to identify statistical measurements that are more dependent on the syntax than on the semantics, the framework may also serve for text analysis in language-dependent applications.	eng
dc.description.version	publishedVersion	eng
dc.identifier.uri	https://doi.org/10.34657/3901
dc.identifier.uri	https://oa.tib.eu/renate/handle/123456789/5272
dc.language.iso	eng	eng
dc.publisher	San Francisco, CA : Public Library of Science (PLoS)	eng
dc.relation.doi	https://doi.org/10.1371/journal.pone.0067310
dc.relation.issn	1932-6203
dc.rights.license	CC BY 3.0 Unported	eng
dc.rights.uri	https://creativecommons.org/licenses/by/3.0/	eng
dc.subject.ddc	530	eng
dc.subject.other	article	eng
dc.subject.other	data analysis	eng
dc.subject.other	information processing	eng
dc.subject.other	linguistics	eng
dc.subject.other	quantitative analysis	eng
dc.subject.other	semantics	eng
dc.subject.other	statistical analysis	eng
dc.subject.other	time series analysis	eng
dc.subject.other	Algorithms	eng
dc.subject.other	Humans	eng
dc.subject.other	Language	eng
dc.subject.other	Models, Statistical	eng
dc.subject.other	Reading	eng
dc.subject.other	Semantics	eng
dc.title	Probing the Statistical Properties of Unknown Texts: Application to the Voynich Manuscript	eng
dc.type	Article
tib.accessRights	openAccess	eng
wgl.contributor	PIK	eng
wgl.subject	Physik	eng
wgl.type	Zeitschriftenartikel	eng

Files

Original bundle

Now showing 1 - 1 of 1

Name:: Amancio et al 2013, Probing the Statistical Properties of Unknown Texts.pdf
Size:: 473.58 KB
Format:: Adobe Portable Document Format
Description:

Download

Collections

Physik