EBookClubs

Read Books & Download eBooks Full Online

EBookClubs

Read Books & Download eBooks Full Online

Book Web Corpus Construction

Download or read book Web Corpus Construction written by Roland Schäfer and published by Springer Nature. This book was released on 2022-05-31 with total page 129 pages. Available in PDF, EPUB and Kindle. Book excerpt: The World Wide Web constitutes the largest existing source of texts written in a great variety of languages. A feasible and sound way of exploiting this data for linguistic research is to compile a static corpus for a given language. There are several adavantages of this approach: (i) Working with such corpora obviates the problems encountered when using Internet search engines in quantitative linguistic research (such as non-transparent ranking algorithms). (ii) Creating a corpus from web data is virtually free. (iii) The size of corpora compiled from the WWW may exceed by several orders of magnitudes the size of language resources offered elsewhere. (iv) The data is locally available to the user, and it can be linguistically post-processed and queried with the tools preferred by her/him. This book addresses the main practical tasks in the creation of web corpora up to giga-token size. Among these tasks are the sampling process (i.e., web crawling) and the usual cleanups including boilerplate removal and removal of duplicated content. Linguistic processing and problems with linguistic processing coming from the different kinds of noise in web corpora are also covered. Finally, the authors show how web corpora can be evaluated and compared to other corpora (such as traditionally compiled corpora). For additional material please visit the companion website: sites.morganclaypool.com/wcc Table of Contents: Preface / Acknowledgments / Web Corpora / Data Collection / Post-Processing / Linguistic Processing / Corpus Evaluation and Comparison / Bibliography / Authors' Biographies

Book Web Corpus Construction

    Book Details:
  • Author : Roland Schäfer
  • Publisher : Morgan & Claypool Publishers
  • Release : 2013-07-01
  • ISBN : 1627053123
  • Pages : 197 pages

Download or read book Web Corpus Construction written by Roland Schäfer and published by Morgan & Claypool Publishers. This book was released on 2013-07-01 with total page 197 pages. Available in PDF, EPUB and Kindle. Book excerpt: The World Wide Web constitutes the largest existing source of texts written in a great variety of languages. A feasible and sound way of exploiting this data for linguistic research is to compile a static corpus for a given language. There are several adavantages of this approach: (i) Working with such corpora obviates the problems encountered when using Internet search engines in quantitative linguistic research (such as non-transparent ranking algorithms). (ii) Creating a corpus from web data is virtually free. (iii) The size of corpora compiled from the WWW may exceed by several orders of magnitudes the size of language resources offered elsewhere. (iv) The data is locally available to the user, and it can be linguistically post-processed and queried with the tools preferred by her/him. This book addresses the main practical tasks in the creation of web corpora up to giga-token size. Among these tasks are the sampling process (i.e., web crawling) and the usual cleanups including boilerplate removal and removal of duplicated content. Linguistic processing and problems with linguistic processing coming from the different kinds of noise in web corpora are also covered. Finally, the authors show how web corpora can be evaluated and compared to other corpora (such as traditionally compiled corpora).

Book Overcoming Challenges in Corpus Construction

Download or read book Overcoming Challenges in Corpus Construction written by Robbie Love and published by Routledge. This book was released on 2020-01-06 with total page 176 pages. Available in PDF, EPUB and Kindle. Book excerpt: This volume offers a critical examination of the construction of the Spoken British National Corpus 2014 (Spoken BNC2014) and points the way forward toward a more informed understanding of corpus linguistic methodology more broadly. The book begins by situating the creation of this second corpus, a compilation of new, publicly-accessible Spoken British English from the 2010s, within the context of the first, created in 1994, talking through the need to balance backward capability and optimal practice for today’s users. Chapters subsequently use the Spoken BNC2014 as a focal point around which to discuss the various considerations taken into account in corpus construction, including design, data collection, transcription, and annotation. The volume concludes by reflecting on the successes and limitations of the project, as well as the broader utility of the corpus in linguistic research, both in current examples and future possibilities. This exciting new contribution to the literature on linguistic methodology is a valuable resource for students and researchers in corpus linguistics, applied linguistics, and English language teaching.

Book Corpus Linguistics and the Web

Download or read book Corpus Linguistics and the Web written by and published by BRILL. This book was released on 2015-07-14 with total page 311 pages. Available in PDF, EPUB and Kindle. Book excerpt: Using the Web as Corpus is one of the recent challenges for corpus linguistics. This volume presents a current state-of-the-arts discussion of the topic. The articles address practical problems such as suitable linguistic search tools for accessing the www, the question of register variation, or they probe into methods for culling data from the web. The book also offers a wide range of case studies, covering morphology, syntax, lexis, as well as synchronic and diachronic variation in English. These case studies make use of the two approaches to the www in corpus linguistics – web-as-corpus and web-for-corpus-building. The case studies demonstrate that web data can provide useful additional evidence for a broad range of research questions.

Book Essential Speech and Language Technology for Dutch

Download or read book Essential Speech and Language Technology for Dutch written by Peter Spyns and published by Springer Science & Business Media. This book was released on 2013-02-26 with total page 414 pages. Available in PDF, EPUB and Kindle. Book excerpt: The book provides an overview of more than a decade of joint R&D efforts in the Low Countries on HLT for Dutch. It not only presents the state of the art of HLT for Dutch in the areas covered, but, even more importantly, a description of the resources (data and tools) for Dutch that have been created are now available for both academia and industry worldwide. The contributions cover many areas of human language technology (for Dutch): corpus collection (including IPR issues) and building (in particular one corpus aiming at a collection of 500M word tokens), lexicology, anaphora resolution, a semantic network, parsing technology, speech recognition, machine translation, text (summaries) generation, web mining, information extraction, and text to speech to name the most important ones. The book also shows how a medium-sized language community (spanning two territories) can create a digital language infrastructure (resources, tools, etc.) as a basis for subsequent R&D. At the same time, it bundles contributions of almost all the HLT research groups in Flanders and the Netherlands, hence offers a view of their recent research activities. Targeted readers are mainly researchers in human language technology, in particular those focusing on Dutch. It concerns researchers active in larger networks such as the CLARIN, META-NET, FLaReNet and participating in conferences such as ACL, EACL, NAACL, COLING, RANLP, CICling, LREC, CLIN and DIR ( both in the Low Countries), InterSpeech, ASRU, ICASSP, ISCA, EUSIPCO, CLEF, TREC, etc. In addition, some chapters are interesting for human language technology policy makers and even for science policy makers in general.

Book Developing Linguistic Corpora

Download or read book Developing Linguistic Corpora written by Martin Wynne and published by Oxbow Books Limited. This book was released on 2005 with total page 100 pages. Available in PDF, EPUB and Kindle. Book excerpt: A linguistic corpus is a collection of texts which have been selected and brought together so that language can be studied on the computer. Today, corpus linguistics offers some of the most powerful new procedures for the analysis of language, and the impact of this dynamic and expanding sub-discipline is making itself felt in many areas of language study. In this volume, a selection of leading experts in various key areas of corpus construction offer advice in a readable and largely non-technical style to help the reader to ensure that their corpus is well designed and fit for the intended purpose. This guide is aimed at those who are at some stage of building a linguistic corpus. Little or no knowledge of corpus linguistics or computational procedures is assumed, although it is hoped that more advanced users will find the guidelines here useful. It is also aimed at those who are not building a corpus, but who need to know something about the issues involved in the design of corpora in order to choose between available resources and to help draw conclusions from their studies.

Book WaCky

    Book Details:
  • Author : Marco Baroni
  • Publisher : Gedit
  • Release : 2006
  • ISBN :
  • Pages : 238 pages

Download or read book WaCky written by Marco Baroni and published by Gedit. This book was released on 2006 with total page 238 pages. Available in PDF, EPUB and Kindle. Book excerpt:

Book Web As Corpus

Download or read book Web As Corpus written by Maristella Gatto and published by A&C Black. This book was released on 2014-02-13 with total page 255 pages. Available in PDF, EPUB and Kindle. Book excerpt: Is the internet a suitable linguistic corpus? How can we use it in corpus techniques? What are the special properties that we need to be aware of? This book answers those questions. The Web is an exponentially increasing source of language and corpus linguistics data. From gigantic static information resources to user-generated Web 2.0 content, the breadth and depth of information available is breathtaking – and bewildering. This book explores the theory and practice of the “web as corpus”. It looks at the most common tools and methods used and features a plethora of examples based on the author's own teaching experience. This book also bridges the gap between studies in computational linguistics, which emphasize technical aspects, and studies in corpus linguistics, which focus on the implications for language theory and use.

Book Construction Grammar and its Application to English

Download or read book Construction Grammar and its Application to English written by Martin Hilpert and published by Edinburgh University Press. This book was released on 2014-03-17 with total page 232 pages. Available in PDF, EPUB and Kindle. Book excerpt: Construction Grammar explains how knowledge of language is organized in speakers' minds. The central and radical claim of Construction Grammar is that linguistic knowledge can be fully described as knowledge of constructions, which are defined as symbolic units that connect a linguistic form with meaning.

Book Corpus Linguistics

Download or read book Corpus Linguistics written by Tony McEnery and published by Cambridge University Press. This book was released on 2011-10-06 with total page pages. Available in PDF, EPUB and Kindle. Book excerpt: Corpus linguistics is the study of language data on a large scale - the computer-aided analysis of very extensive collections of transcribed utterances or written texts. This textbook outlines the basic methods of corpus linguistics, explains how the discipline of corpus linguistics developed and surveys the major approaches to the use of corpus data. It uses a broad range of examples to show how corpus data has led to methodological and theoretical innovation in linguistics in general. Clear and detailed explanations lay out the key issues of method and theory in contemporary corpus linguistics. A structured and coherent narrative links the historical development of the field to current topics in 'mainstream' linguistics. Practical tasks and questions for discussion at the end of each chapter encourage students to test their understanding of what they have read and an extensive glossary provides easy access to definitions of technical terms used in the text.

Book Building and Exploring Web Corpora  WAC3   2007

Download or read book Building and Exploring Web Corpora WAC3 2007 written by Cédrick Fairon and published by Presses univ. de Louvain. This book was released on 2007 with total page 186 pages. Available in PDF, EPUB and Kindle. Book excerpt: WAC More and more people are using Web data for linguistic and NLP research. The Web as Corpusworkshop (WAC) provides a venue for exploring how we can use it effectively and the advancementsto which this could lead.This book is a collection of the talks presented at the 3 rd WAC in Louvain-la-Neuve (Belgium).The focus is on the description of Web corpus collection projects, the exploration of Web datacharacteristics from a linguistics/NLP perspective, and on the use of crawled Web data for NLPpurposes. CLEANEVAL Any use of Web data requires that it be cleaned in order to get rid of unwanted material including,for example, HTML markup, navigation bars, advertisements. To date there has been no sharingof resources or expertise in this particular domain and the cleaning has often been done minimally.Cleaneval was an exercise aimed at promoting collaboration and improving our understandingof the issues. Results and perspectives are presented in this book.

Book Corpus linguistics

Download or read book Corpus linguistics written by Stefanowitsch, Anatol and published by Language Science Press. This book was released on 1996 with total page 510 pages. Available in PDF, EPUB and Kindle. Book excerpt: Corpora are used widely in linguistics, but not always wisely. This book attempts to frame corpus linguistics systematically as a variant of the observational method. The first part introduces the reader to the general methodological discussions surrounding corpus data as well as the practice of doing corpus linguistics, including issues such as the scientific research cycle, research design, extraction of corpus data and statistical evaluation. The second part consists of a number of case studies from the main areas of corpus linguistics (lexical associations, morphology, grammar, text and metaphor), surveying the range of issues studied in corpus linguistics while at the same time showing how they fit into the methodology outlined in the first part.

Book Text  Speech and Dialogue

Download or read book Text Speech and Dialogue written by Petr Sojka and published by Springer. This book was released on 2014-09-01 with total page 623 pages. Available in PDF, EPUB and Kindle. Book excerpt: This book constitutes the refereed proceedings of the 17th International Conference on Text, Speech and Dialogue, TSD 2013, held in Brno, Czech Republic, in September 2014. The 70 papers presented together with 3 invited papers were carefully reviewed and selected from 143 submissions. They focus on topics such as corpora and language resources; speech recognition; tagging, classification and parsing of text and speech; speech and spoken language generation; semantic processing of text and speech; integrating applications of text and speech processing; automatic dialogue systems; as well as multimodal techniques and modelling.

Book The Handbook of Asian Englishes

Download or read book The Handbook of Asian Englishes written by Kingsley Bolton and published by John Wiley & Sons. This book was released on 2020-09-14 with total page 928 pages. Available in PDF, EPUB and Kindle. Book excerpt: The first volume of its kind, focusing on the sociolinguistic and socio-political issues surrounding Asian Englishes The Handbook of Asian Englishes provides wide-ranging coverage of the historical and cultural context, contemporary dynamics, and linguistic features of English in use throughout the Asian region. This first-of-its-kind volume offers a wide-ranging exploration of the English language throughout nations in South Asia, Southeast Asia, and East Asia. Contributions by a team of internationally-recognized linguists and scholars of Asian Englishes and Asian languages survey existing works and review new and emerging areas of research in the field. Edited by internationally renowned scholars in the field and structured in four parts, this Handbook explores the status and functions of English in the educational institutions, legal systems, media, popular cultures, and religions of diverse Asian societies. In addition to examining nation-specific topics, this comprehensive volume presents articles exploring pan-Asian issues such as English in Asian schools and universities, English and language policies in the Asian region, and the statistics of English across Asia. Up-to-date research addresses the impact of English as an Asian lingua franca, globalization and Asian Englishes, the dynamics of multilingualism, and more. Examines linguistic history, contemporary linguistic issues, and English in the Outer and Expanding Circles of Asia Focuses on the rapidly-growing complexities of English throughout Asia Includes reviews of the new frontiers of research in Asian Englishes, including the impact of globalization and popular culture Presents an innovative survey of Asian Englishes in one comprehensive volume Serving as an important contribution to fields such as contact linguistics, World Englishes, sociolinguistics, and Asian language studies, The Handbook of Asian Englishes is an invaluable reference resource for undergraduate and graduate students, researchers, and instructors across these areas. Winner of the 2021 PROSE Humanities Category for Language & Linguistics

Book Ten Lectures on Diachronic Construction Grammar

Download or read book Ten Lectures on Diachronic Construction Grammar written by Martin Hilpert and published by BRILL. This book was released on 2021-09-13 with total page 291 pages. Available in PDF, EPUB and Kindle. Book excerpt: In this book, Martin Hilpert lays out how Construction Grammar can be applied to the study of language change. In a series of ten lectures on Diachronic Construction Grammar, the book presents the theoretical foundations, open questions, and methodological approaches that inform the constructional analysis of diachronic processes in language. The lectures address issues such as constructional networks, competition between constructions, shifts in collocational preferences, and differentiation and attraction in constructional change. The book features analyses that utilize modern corpus-linguistic methodologies and that draw on current theoretical discussions in usage-based linguistics. It is relevant for researchers and students in cognitive linguistics, corpus linguistics, and historical linguistics.

Book Contemporary Corpus Linguistics

Download or read book Contemporary Corpus Linguistics written by Paul Baker and published by A&C Black. This book was released on 2012-03-15 with total page 370 pages. Available in PDF, EPUB and Kindle. Book excerpt: Acts as a one-volume resource, providing an introduction to every aspect of corpus linguistics as it is being used at the moment.

Book Corpus Linguistics and the Description ofEnglish

Download or read book Corpus Linguistics and the Description ofEnglish written by Hans Lindquist and published by Edinburgh University Press. This book was released on 2009-12-07 with total page 241 pages. Available in PDF, EPUB and Kindle. Book excerpt: A lively hands-on introduction to the use ofelectronic corpora in the description and analysis of English, this bookprovides an ideal introduction for university students of English at theintermediate level. Students planning papers, dissertations or theses willfind the book a particularly valuable guide.After introducing corpora andthe rationale and basic methodology of corpus linguistics, the authorpresents a number of case studies providing new insights into vocabulary,collocations, phraseology, metaphor and metonymy, syntactic structures, maleand female language, and language change. In a final chapter it is shown howthe web can be used as a source for linguistic investigations. Each chapterhas study questions, exercises and suggestions for further reading.Studentswill benefit from the book's*Clear language and structure *Well-definedterminology *Step-by-step instructions *Generous, up-to-date exemplificationfrom different varieties of English around the world *Accompanying web-pagewith exercises and updated information about freely accessiblecorpora.