Interviews With Researchers – Europeana Newspapers http://www.europeana-newspapers.eu A Gateway to European Newspapers Online Mon, 13 Feb 2017 14:07:19 +0000 en-US hourly 1 https://wordpress.org/?v=6.2.9 Q&A with newspaper researchers: Sophie Kurkdjian http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-sophie-kurkdjian/ http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-sophie-kurkdjian/#comments Fri, 20 Mar 2015 16:12:36 +0000 http://www.europeana-newspapers.eu/?p=3665 Continue reading Q&A with newspaper researchers: Sophie Kurkdjian]]> Sophie Kurkdjian
Photo courtesy of Sophie Kurkdjian

Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering?

This month we’re interviewing Sophie Kurkdjian, researcher at the IDHES – Institutions et dynamiques historiques de l’économie et de la société of the Université Paris 1 in France.

Can you briefly describe yourself: your background and the research you’ve done using historic newspapers?

I hold a Master degree in contemporary history on the history of journalists in the women’s magazines, based on the study of titles published between 1945 and 2000: mainly Elle and Marie-Claire and daily newspapers such as Le Monde and Le Figaro.

Vogue, October 1927
Vogue, October 1927

I have also written my Doctoral thesis at the University Paris 1 Panthéon-Sorbonne on the life of two news publishers at the beginning of the 20th century – Lucien Vogel and Michel Brunhoff – with researches based on illustrated magazines from the first half of the 20th century, women’s magazines (Gazette du bon ton, Jardin des modes, Vogue français, Vogue), illustrated press (Vu, Regards, BIZ), daily press and weekly literary-political press (Le Petit Journal, Marianne, Messidor).

 

Your work highlights the importance of newspapers as an information source. How did you discover the link between your research topic and the information recorded by the press?

The press was more than a source of information for me. It was the core topic of both my studies themselves. The link between my research and the press was indeed very important. The central part of my work involved analysing the information in the chosen newspapers in terms of display (layout, title, colours…) and content (history, social policy, and cultural policy).

Was this the first time that you had used newspapers as sources of information for research and did it change the way that you perceive newspapers?

During my research, I faced for the very first time the fact that I was using the press both as a main source of information and as an object of study. While writing my Doctoral thesis I have used newspapers in a more throrough way and it has boosted my interest for this type of material in the context of the new posibilities offered by digitisation, which enhances the richness of such a source of information, not to say the interest to work on it.

How would you compare newspapers to other sources of information such as books and journals?

Vogue, May 1930
Vogue, May 1930

Newspapers constitue a very rich source of information and first of all allow to work on diverse aspects like format, content, readership, influence on other titles, and reading habits. They offer various entries and can be approached from different points of view (political, economical, or social). One major aspect that can be found in newspapers (and which renders them more interesting and makes them unique in comparison with other sources) is the closeness to the reader’s daily life, to a society at a particular time and place.

They constitute mirrors, seismographs of socio-cultural changes. They reveal social changes such as new consumption and reading habits. They show how readers got used progressivly to reading images: to understand an event through a picture. The study of format changes shows more than any other source of information the evolution of collective representations and mentalities, but also gaps between certain assumptions and the readers “real” practices.

What specific types of information did the newspapers contain that you found valuable?

Social, cultural and political data, the social mood a specific moment (in the readers mail for example) are all valuable information.

In terms of your workflow, did you use digital or paper copies of newspapers and what kind of techniques did you use (e.g. simple keyword search, text mining)?

I used simultinously both digital and print sources. When it comes down to technical aspects, I used the keyword search functionality to access information one would not be able to find in a print version or that would consume much of the time for investigation.

How was your research affected by the format of newspapers that you used, in both a positive and in a negative sense?

To hold in hand a print copy reveals key indications regarding the size, the color of the cover for example, how a format in particular illustrates a change, the importance (or the lack) of advertising in a magazine.

Vogue, March 1940
Vogue, March 1940

The digital version allows a more general, comparative approach and, as a consequence, to work on a much more precise chronology, to assess with a greater accuracy format changes for instance or the nature of the social debates conveyed by the newspaper, to deconstruct certain ideas through studying more precisely various titles. The negative point lies in the difficulty to set a limit to our field of investigation, to restrain our researches.

 

If you could choose today between using a digital or a paper archive, which would you choose and why?

I would favor the digital one for practical reasons and the time gain. It allows for a more global vision of a newspaper and even a comparative one. The possibility to select texts and to use the keyword search increases search possibilities.  However, for certain types of newspapers such as illustrated ones, the print version still matters in order to understand the layout, graphism, colours, the rendering, and especially to understand how we entered in the era of the visual society where news had to be shown rather than described.

Looking forward, how would you improve access to historic newspapers? Are there specific tools that need to be provided, or needs that should be met by libraries and digital archives?

Rather than creating new tools, I would prefer specifically targeted improvements. Regarding press magazines for instance, one should take care of not putting all summaries (or all the commercials) for a month/year at the end of the document/volume of the digital copy to be found in Gallica.  Advertisements should be associated with each summary and each issue so that their study can be improved.

What potential do you see for a pan-European archive such as the one being built by Europeana Newspapers? Could you, for example, extend your thesis by having access to newspapers from across Europe via a single website?

The major interest of such a project will be to understand, in a comparative approach, how newspapers from different countries are connected in terms of influence or kinship. For instance, British and German illustrated periodicals influenced not only the French but also the American press at the turn of the 20th century. Europeana Newspapers is an ambitious project which will appear as a valuable tool for researchers who encounter difficulties while working on foreign newspapers.

Sophie Kurkdjian, many thanks for this insightful interview.

]]>
http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-sophie-kurkdjian/feed/ 2
Q&A with newspaper researchers: Timo Honkela http://www.europeana-newspapers.eu/timo-honkela/ http://www.europeana-newspapers.eu/timo-honkela/#respond Thu, 30 Oct 2014 13:43:29 +0000 http://www.europeana-newspapers.eu/?p=3131 Continue reading Q&A with newspaper researchers: Timo Honkela]]>

 Within the Europeana Newspapers Project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering? This month we’re interviewing Timo Honkela, professor at University of Helsinki, Department of Modern Languages, and National Library of Finland, Centre for Preservation and Digitisation.

 

Photograph from Timo Honkela made by Päivi Piispa, National Library Finland

Can you briefly describe yourself: your background and the research you’ve done using historic newspapers?

My name is Timo Honkela and from the beginning of 2014, I have served as a professor at University of Helsinki, Department of Modern Languages, and National Library of Finland, Centre for Preservation and Digitisation. In 1980s, I worked as a researcher in a project that was developing a natural language database interface for Finnish. Since then, the research has been around natural language processing, machine learning, text mining, cognitive modeling and related areas

 

Your work highlights the importance of newspapers as an information source. How did you discover the link between your research topic and the information recorded by the press?

Newspapers are a natural information source in our research. First of all, the  Centre for Preservation and Digitisation in Mikkeli, Finland, is respobsible for storing historical and modern newspapers in digital form. At the same time, it is a major source regarding Finnish language, history and culture.

 

Was this the first time that you had used newspapers as a source of information for research, and did it change the way that you perceive newspapers? In other words, were you surprised by what you could research using newspapers?

In earlier research, newspapers or news feed was used to test the Websom system. This system for organizing text documents automatically into an organized map has been a forerunner in the area of information visualization and widely cited in scientific literature. In that purpose, newspapers were dealt with rather superficially. In our current research, we consider the contents in a more detailed manner.

 

How would you compare newspapers to other sources of information such as books and journals? Are there certain aspects of newspapers that just can’t be replicated anywhere else?

Newspapers provide an in-depth view into the functioning of a whole society. The details regarding people, places and events are often such that cannot be found elsewhere.

 

In terms of your work process, did you use digital or paper copies of newspapers and what kind of techniques did you use (eg. simple keyword search, text mining)?

Our work is essentially text mining. Therefore, it is important that the newpapers are available in digital form. Regarding the historical newspapers, one problem is that the error rate in the output of the Optical Character Recognition is often rather high and sometimes not useful at all. Therefore, one of the first tasks is to device automatical postprocessing that corrects some proportion of the errors. Various approaches can be used including language technology and statistical machine learning. The newspaper corpus contains mainly texts in Swedish and Finnish. As there are millions of pages, manual revision work is not possible. Finnish introduces challenges due to its complex morphology. Every Finnish noun has about 2,000 inflectional word forms and every verb more than 10,000.  Therefore, one cannot simply list Finnish words and compare the OCR output with this kind of a list. Regarding complexities of Finnish, collaboration is conducted with Dr. Krister Lindén who directs FIN-CLARIN consortium.

 

How do you use machine learning?

We apply machine learning methods and techniques in various ways. One basic approach is to build a language model. This model can predict the next letter in a word or the next word in a sentence based on a probabilistic model. We also conducted conceptual modeling based on unsupervised learning algorithms such as the self-organizing map and independent component analysis. Machine learning can also device various other tasks such as term extraction and sentiment analysis.

 

Looking forward, how would you improve access to historic newspapers? Are there specific tools that need to be provided, or needs that should be met by libraries and digital archives?

Actually, this question relates nicely to the earlier discussion on improving the quality of scanned texts and finding means to text mine the contents of the articles. One could, in addition, build new kinds of views into the collections using  visual text mining. The Websom project developed suitable techniques for this already in the 1990s. On the other hand, these kind of sophisticated techniques have become feasible only recently thanks to better computational resources. Another interesting area is the use of multilingual language technologies. In the META-NET Network of Excellence, we were involved in taking this area further. It seems evident that the use of machine translation and cross-lingual information retrieval will increase.

 

What potential do you see for a pan-European archive such as the one being built by Europeana Newspapers? Could you, for example, extend your thesis by having access to newspapers from across Europe via a single website?

A pan-European archive and European collaboration in this area will be very important. In our institution, we have been discussing a platform for historical and societal research based on European newspaper collections. New text mining and machine translation technologies would enable research in which important historical developments and societal questions could be analyzed in an unforseen manner. More specifically, we have developed a method called Grounded Intersubjective Concept Analysis (GICA) that can be used to analyze similarities and differences between concepts and conceptual systems among people and gorups of people. This and related methods can be used to analyze corpora of hundreds of millions of European newspaper pages to find out how different conceptions and perspectives into important historical events and societal themes have developed over time.

Another interview with Timo Honkela on ” Making the most of digital materials” is available here on the website of the National Library of Finland.

 

 

]]>
http://www.europeana-newspapers.eu/timo-honkela/feed/ 0
Q&A with newspaper researchers: Antal van den Bosch http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-antal-van-den-bosch/ http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-antal-van-den-bosch/#respond Thu, 04 Sep 2014 08:27:46 +0000 http://www.europeana-newspapers.eu/?p=3028 Continue reading Q&A with newspaper researchers: Antal van den Bosch]]>

Within the Europeana Newspapers Project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering? This month we’re interviewing Antal van den Bosch, full professor at Centre for Language Studies and Communication and Information Sciences at the Radboud University Nijmegen in the Netherlands.

AntalvandenBosch2

Could you briefly describe your background and your main fields of research?

My background is in computational linguistics. I received a Master’s degree in computational linguistics in 1992 from Tilburg University at the faculty of Arts. Afterwards I went to Maastricht University also in the Netherlands where I did my PhD research at the computer science department. I continued working on computational linguistics especially on the application of machine learning, e.g. text to speech synthesis and text to text processing.

Over the years I developed interest in hard challenges in the area of machine learning of natural language, e.g. in word sense disambiguation, machine translation, and language modeling. It is a busy field but still a lot to be done.

Historical texts and archived materials have always been interesting for linguists. Data gathering has also been going on for quite some time, e.g. collections of newspaper archives, literary works, or folk tales. We, as computational linguists, generally analyze those materials on specific linguistic levels and generate automatic classifications for complete texts or cluster texts within the archive.

 

You used historic newspapers as a source in a project to find strikes that actually never happened – Could you please expand a bit more on the project and how it was related to digitized historic newspapers?

Actually, there were two projects. The first started five years ago, finished last year and was called HiTiME, which is an abbreviation of “Historical Timeline Mining and Extraction”. The project was carried out at Tilburg University and was about mining historical texts and text analysis. Our partner institute, the International Institute for Social History in Amsterdam, has a lot of databases and also texts covering all matters relating to the history of the social movement. One database was about strikes in the Netherlands; another contained a list of all labor unions that ever existed in the Netherlands. They also had terminology lists and authority files, basically structured pieces of information. They needed us to link them up and semantically connect them. Our question to them was: What do you want to do with that information? We asked several historians which kind of research they do, which kind of research they would like to do and so on. And they gave us a lot of hints. One of the things they often said was that they were interested in the dynamics of things. How do strikes evolve over time, what are the points of no return that lead to either nothing or to a strike? It is like at a crossroads. You rarely see an accident but often situations occur that almost go wrong. For researchers, strikes that did not happen are as interesting as strikes that did happen.

We linked this question to the digital collection team of the National Library of the Netherlands in The Hague, who were close to finishing their big newspaper digitization project.

We took the database that Sjaak van der Velden has been working on at the International Institute of Social History. It is a database of all the strikes that ever occurred in the Netherlands with information on where each strike took place, when it started, when and how it ended and so on. We took every item in the database and turned it into a query that said: find me articles that contain words like staken which means to strike in Dutch or staakten – the past tense – plus words from the database entry indicating location or factory name. We also restricted the dates in the query to a week before and three days after the date from the database. For each query we got a lot of articles. So we roughly linked articles to strikes. We focused on a couple of periods that were kind of hot according to historians. From these sample periods we got a lot of articles that were talking about strikes. If this produces fairly reliable data, which it did, you can build a classifier similar to a spam filter. You can then filter articles for strike threats for example by feeding the classifier with positive and negative examples, where positive examples are articles that are written before a strike. That filter can be applied to the complete archive. The result is a selection of articles that point at a particular moment of social unrest where a strike either was going to happen or not. This is where we stopped because it then becomes the task of the historian to work with the data, classify it, and look for patterns.

The gain is that we are saving the historian a lot of time. We prioritized the search. The historians that followed us said that there is no manual alternative to this type of very fast method.

 

Which were the main challenges you faced when working on these projects?

There were a few challenges: One was the question: How do you do scale up? – At the beginning the National Library sent us hard disk drives with their files from the archive, because it isn’t possible to apply the filter on the whole archive over the internet. The network can’t handle it. Then there is the issue of OCR errors. The National Library did a great job but OCR hasn’t reached a sufficient quality level to guarantee precise searching and retrieval results at the moment, especially with older newspapers. Due to their paper quality for example, OCR can be really bad. Although fuzzy search is available as well for the archive, one person within our group, Martin Reynaert, developed new post-correction tools for OCR. So work is carried out to solve this problem but it isn’t at the moment.

 

Was this the first time that you had used historical newspapers as a source and did the experiences you gathered from working with that source influence your later work? 

We worked with digitally-born newspaper archives in the 1990s and 2000s by asking newspaper companies. There is a lot of applied work on current news, a lot on commercial and social media. The work with the National Library was our first encounter with historical newspapers; they had just released their newspaper archive. Now it has a lot of users but we were one of the first.

 

If one thinks of researchers working with historical newspapers one frequently thinks of historians or literary scholars. Your background is quite different. Can other disciplines use digitized newspapers for their research as well?

Speaking for my own field, we regard texts rather agnostically. Text is data for us. But if I may generalize a little, linguists have a big interest in such corpora. They study variation and change in language over time and geographical spaces. Newspapers give a fairly good idea about what was common to talk about in a certain region or time. I know some studies based on newspapers from Dutch colonies such as Surinam or Indonesia. The language of these newspapers shows interesting differences to standard Dutch newspapers. Some words came into use in the colonies and diffused into general Dutch in interesting periods both from a linguistic as well from a historical point of view.

 

Is the gathered data collection interesting for completely different areas of research as well, e.g. as testing grounds for software tools, which are then utilized in another area?

You can find examples of that. Scaling is always an issue in technology but there are other archives available as well. The internet itself is a much bigger archive. Social media alone has produced a lot of data. In the Netherlands alone about 7 million tweets are posted every day and this since about 2010, when it started to be a mainstream type of social media. This is creating an amazing amount of data.

But for the issue of searching and indexing, having those archives available I think is really great as a scientific challenge, because of its noisiness. How do you reliably link canonical, present-day keywords to words with historical spellings, with and without OCR errors? How do you access this enormous amount of data which is noisier than average internet data?

 

You mentioned data mining. As a rather new field of research this area becomes more and more an issue. Could you maybe expand a bit more on the concepts and research interests behind that term?

Data mining is obviously more general than text mining. It is about exploring and discovering patterns that haven’t been linked so far in any kind of data. For example, newspapers reflect the economy or economic events and provide economic data in great detail. Often in retrospective research, e.g. socio-economic historical research, official publications have been used (e.g. company press releases and official publications of central statistics bureaus) whereas a newspaper has more to offer. In the official documents you might have four data points a year per company, while newspapers talk maybe a hundred times a year about the same company. For economic historians this is an interesting extra layer not present in databases so far.

 

You are also using Twitter as a source for research. What are the similarities and differences between an historical newspapers and Twitter – its sort of contemporary counterpart?

I think that historical newspapers are more subjective than modern newspapers. In the Netherlands each had a very clear profile as e.g. social, liberal, catholic. They were rather antagonistic against one or two of the other groups. You can see that very clearly. And they were written by journalists who were not trained as they are nowadays, where they are academics or have had professional training. This profession has definitively grown. What you see in Twitter, in contrast, is that obviously most users are not professional writers. You see subjectivity in Twitter in a higher degree than in newspapers. Another difference in Twitter is that tweets come rather close to speech. This is different from newspapers in all ages as it always has been a written genre and intentionally not too colloquial – which Twitter is to a large extent.

 

What potential do you see for a pan-European digital collection such as the one being built by Europeana Newspapers?

The big picture should of course be on a European and global scale. Social processes such as how and why labor unions organize cannot be fully understood unless you go beyond the borders of a particular country. Social movements in the Netherlands can’t be seen without European movements. The Netherlands were in some cases trendsetting but usually followed other examples.

Studies that have been done on a smaller, monolingual and single-country scale should be lifted to a multilingual level. We did a second project after HiTiME with partners from the UK and US. It was funded by Digging into Data. We basically extended the study to English and got access to the archive of the New York Times. In essence we did the same analysis of strikes and news reports on strikes and reproduced the study. This is an example what you should do: get multiple perspectives. The New York Times did not only cover major strikes in the U.S., especially on the East Coast and New York, but also strikes in Europe. The American perspective, on a European strike is obviously framed within the international relations of the countries. For example when there was a strike of coal miners, the Americans reported on the consequences for import and export.

 

Looking into the future, how would you improve access to historic newspapers? Are there specific tools that need to be provided, or needs that should be met by libraries and digital archives?

A good question and also not easy to answer. It is great that they are available. One issue I was pointed at by the National Library in The Hague is that they are rescanning things. It is great if the quality of the data improves. But they may not be providing access to all versions of newspapers they are scanning. The reproducibility of our experiments therefore becomes a problem and version control an issue. We did our study but if the data is not in our place, and it shouldn’t be because of copyright issues, we can only refer to the data at the National Library. If they change even small things, the results of applying our methods to the Library’s current data will be different. This goes against a basic scientific principle that it should be possible to replicate results. It is not enough to publish our software; it needs to be the same for data. Libraries currently have trouble guaranteeing that because they may have trouble providing enough storage and retrieval or backup space.

Besides more storage and good backup schemes, there is a need for version control. We develop software that is dynamic but you need to know which software is used and it must be possible to roll back to older versions; the same goes for data. This is where all this v. 1.0, v. 2.0 and so on is coming from. So versioning in computer science is kind of an art but also an obligation. The same should go for digital collections as well.

Mr. Van den Bosch, many thanks for this insightful interview.

 

]]>
http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-antal-van-den-bosch/feed/ 0
Q&A with newspaper researchers: Amélia Sanz Cabrerizo http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-amelia-sanz-cabrerizo/ http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-amelia-sanz-cabrerizo/#respond Thu, 17 Jul 2014 07:33:02 +0000 http://www.europeana-newspapers.eu/?p=2820 Continue reading Q&A with newspaper researchers: Amélia Sanz Cabrerizo]]> amelia-del-rosario-1Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering?

This month we’re interviewing Amélia Del Rosario Sanz Cabrerizo, a researcher and professor at Universidad Complutense de Madrid. She teaches comparative literature and cyber culture (general understanding of challenges of cyber developments; digital means, transformation of mentalities in daily lives, etc.), as well as an introduction to using databases for undergraduates.

1. Can you briefly describe yourself: your background and the research you’ve done using historic newspapers?

I have worked on Comparative Literature from my PHD up to now, mainly Spanish and French Cultural Relationships. At the very beginning I worked on influences and reception, but I have been moving towards the notions of circulation and transfer, I mean from a national (or nationalist approach) to a more local and global one. At that time, I used to look for one type of item in a single magazine: for example, serialised French novels with German characters in one concrete French Journal in 1917. So I was very familiar with research on just one newspaper, looking for texts mainly.

2. Your work highlights the importance of newspapers as an information source. How did you discover the link between your research topic and the information recorded by the press?

From the 18th century up to now, newspapers have been a primary source for researchers working on how public opinion is built. I read Marc Angenot’s works, particularly 1889 as an example of systematic analysis from newspapers.

Curiously, journalese — the language used in newspapers — is  considered a marginal corpus for linguistics, but not for us, as researchers on cultural phenomena: it is one of our main corpus:

  • Because it is highly edited, so elaborated, a discursive construction;
  • With very specific restrictions, I mean subject to medium;
  • It is a plural discourse since it involves a lot of voices and not just influenced by one single author, but by a huge amount of them.

3. Was this the first time that you had used newspapers as a source of information for research, and did it change the way that you perceive newspapers? In other words, were you surprised by what you could research using newspapers?

For my PHD, I worked on last 17th century Spanish romans reception, when newspapers were not so many and not so important, but when I began to work on 18th and 19th century, and after reading Marc Angenot as I said, It was obvious that I had to use newspapers. The problem was which ones (because it was a very time-consuming task) and where (because I had to go to the National Library in Paris or in Madrid so it was a very hard and expensive task).

4. How would you compare newspapers to other sources of information such as books and journals? Are there certain aspects of newspapers that just can’t be replicated anywhere else?

We know that newspapers are the most important format and the most important formula for circulation of ideas, not only in a local context but in a much broader one. If you are interested in circulation, you have to work with newspapers. An example: in the Spanish context, some newspapers are very, very important  because they were read in South America as much as in Spain. Mapping this circulation should be a priority in the future, not before massive digitisation processes and virtual libraries.

5. What specific types of information did the newspapers contain that you found valuable, and why were these important for you?

At present, I am using digitised newspapers available in virtual libraries to look for allusions, quotations, publications, translations to some women writers. I am trying to prove that these women writers were read all over Europe but they were banished by critics in the 19th and 20th centuries. I can provide quantitative analysis to prove the importance of some women in their days, or the insignificance of some other women who were consecrated a long time after.

7. In terms of your work process, did you use digital or paper copies of newspapers and what kind of techniques did you use (eg. simple keyword search, text mining)?

I am using digital copies exclusively for research on term extractions from the magazines and newspapers  for a very specific period. We use a simple keyword search, even if I have some problems of spelling due to the OCR systems (separation of letters, more towards the end of 18th century or early 19th century). It is never an exhaustive analysis but it is a very interesting experience for students with whom I am working in the frame of the two main Spanish Digital Newspapers libraries: Hemeroteca Digital at the National Library, and the Virtual Library of Historical Newspapers.

8. How was your research affected by the format of newspapers that you used, in both a positive and a negative sense?

Browsing through paper copies might accidently lead you to other information (eg. notes on the side of the newspaper, stories in an edition that is located in close physical proximity to the issue you originally were looking through), while being able to text mine might allow you to make connections across a far larger corpus of work than would be possible with paper copies.

You cannot browse through paper copies from page to page, from newspaper to newspaper, looking for some allusions to one concrete author or work.  Of course you will find hundreds of papers in books and reviews and hundreds of communications on the presence of one single author  or one very specific question in a certain newspaper, but is it representative enough? For the old scientific paradigm, perhaps; but not for the new one, because our experimental field is larger and larger. To speak as Popper does: nowadays it is a question of falsifiability.

9. If you could choose today between using a digital or a paper archive, which would you choose and why?

Of course, I would choose a digital archive, because:

  • I am saving time and costs;
  • I can handle a huge amount of data understandable by machines.

I’d rather a paper archive if and only if the newspapers are not digitised and if (and only if) I consider they are very important for my research. This raises a central question: it is a fact that 30% of heritage collections will not be digitised (I am quoting the 2nd Survey  elaborated and published by Enumerate). Why? What about them? Which are the criteria used to select the archive (and I am quoting Michel Foucault) , to create the new digital archives? Who is making these (informed) decisions?

Researchers will not trust fully digital libraries because 30% of the content will not be digitised.

10. Looking forward, how would you improve access to historic newspapers? Are there specific tools that need to be provided, or needs that should be met by libraries and digital archives?

First of all, I advocate for an organisational change; a more user-oriented scope in order to bridge the gap between providers and scholars. Some initiatives such as this one are great, but also:

  • Please make the environment more friendly: where is the information desk? Where is the contact in Europeana? Are there any Advisory Boards for scholars, researchers, students? Any menu with a clear list of services to help/teach us?
  • A user-friendly means of searching across multiple European newspapers, including a simple annotation tool so that users can add comments or clarifications to items and data.
  • A training section with services for research and learning communities, tools available for researchers and learners, even a selection of good practices.
  • A kind of virtual Lab  for researchers (such as in the British Library in the UK or in the Netherlands).
  • Unrestricted access, rather than more money for tools, and ways to feed your raw data into my database.

In addition, you have to improve the quality of digital OCR. It could also be useful to have a centralised index of newspapers for every digital newspaper library which has aggregated data and material to Europeana: How many newspapers have been digitised and, the most important point, what were the criteria to choose these selected newspapers? Why these ones and not those ones?

Without this information, a real pan-European comparison is not possible. We have to be sure that our corpus (the newspapers available, the selected newspapers) is representative enough and balanced enough (with regard to a particular target designated by the researcher) to be analysed automatically.

Finally, it is a time to move from conservation to conversation. Please have more conversations such as this one, a nice initiative.

What potential do you see for a pan-European archive such as the one being built by Europeana Newspapers? Could you, for example, extend your thesis by having access to newspapers from across Europe via a single website?

Having access to newspapers from across Europe means we are able to go beyond binarist comparisons (and the logic of antagonisms) in favour of a more plural scope: marginal, peripherical and central; small and big cultures. It allows us to look for circulation, not only for origins: to study routes rather than roots, to work on what we call transliteratures.

Also, the possibility of managing a great amount of data with machines could change completely our perception of the importance of authors, works, concepts: Mme Cottin vs Mme de Grafigny. We can quantify their presence all over Europe and make a qualitative analysis.

]]>
http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-amelia-sanz-cabrerizo/feed/ 0
Q&A with newspaper researchers: Kārlis Vērdiņš http://www.europeana-newspapers.eu/karlis-verdins/ http://www.europeana-newspapers.eu/karlis-verdins/#respond Tue, 29 Apr 2014 06:48:33 +0000 http://www.europeana-newspapers.eu/?p=2733 Continue reading Q&A with newspaper researchers: Kārlis Vērdiņš]]> Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering?

Kārlis Vērdiņš. Image by Ave Maria Mõistlik and licensed under CC BY 3.0
Kārlis Vērdiņš. Image by Ave Maria Mõistlik and licensed under CC BY 3.0

This month we’re interviewing Kārlis Vērdiņš, a researcher at the Institute of Literature, Folklore and Art at the University of Latvia. He has written a formidable monograph about free verse in poetry (“The Bastard Form”, 2011) and also is a renowned Latvian poet and translator.

What is your area of research?
The main field of my research is the history of Latvian literature. I am interested also in queer theory and gender studies. Currently I’m researching topics of gender and sexuality in Latvia at the turn of 20th century. Together with three of my colleagues we are working on a book about this period. Its working title is Translating Cultures, referring to the fact that during this period Latvian culture was borrowing a lot from other cultures, adopting or translating customs and ideas in its own fashion.

I know that your current research also involves periodicals.
Yes, in my recent work the possibilities to use digitised periodicals from the National Library of Latvia (www.periodika.lv) were crucial. This resource makes working with periodicals of late 19th century and the first half of 20th century much more productive. First of all, digitised newspapers are easier to access, it can be done from my home computer. I have spent many late nights in the periodical’s portal, and there are entirely new possibilities to explore usage of various terms. Due to the wear and tear of the originals, sometimes the digitised material is not perfect, but it is still useful. The texts are easy transformable into digital format, a Word document etc. And the use of keywords make easy to work with it.

Periodika, the digital newspapers portal of the National Library of Latvia.
Periodika, the digital newspapers portal of the National Library of Latvia.

What is special about periodicals, why are they particularly interesting?
Sometimes we imagine that the topics and problems we encounter today are entirely new. Newspapers and magazines prove that it is not quite so. This is very much the case also with the sexuality and homosexuality in particular. The topic of homosexuality is virtually non-existent in Latvian books of the late 19th and early 20th century but it surfaces in newspapers.

An issue of the Preciniece/Svaha newspaper, from the Periodika newspaper portal.
An issue of the Preciniece/Svaha newspaper, from the Periodika newspaper portal.

What did you find?
I’ve found many interesting facts and publications concerning the history of homosexuals in Latvia. For example, there was a Latvian actor Roberts Tautmīlis Bērziņš who was constantly mocked for his femininity in the performance reviews and gossip columns. However, there was one performance where his acting was widely praised – when he impersonated a student who for the sake of a prank pretends to be the rich aunt from Brasil.

Another exciting discovery was finding the only pro-gay newspaper in the 1920s – a humble and shortlived bilingual edition of marriage ads and short love stories. Preciniece/Svaha (The Matchmaker) published, in its last issues, articles on gay themes by an unknown Russian author with a penname Siluet. For sure, I wouldn’t come across this information if I were only looking through physical newspapers, the search possibilities in digitised material were crucial.

What are the main problems when working with digitised material?
The problem is that the resources are not complete yet. At times there has been a confusion regarding older and newer versions of the resources – some editions appeared in one version and seemed lost in another, and so on. Also you cannot obtain a complete picture of the epoch if, for example, the exile publications are very thoroughly represented while some of the most important local titles are still absent. On the other hand, the collections of newspapers have already been successful in merging sets of periodicals which are partially incomplete in several libraries. The problem of incompleteness is probably less acute in a small country like Latvia. In the long term, it will be easier for us to digitise virtually everything, even the materials of comparatively secondary importance (I certainly would like to see it happen), while large cultures probably will never digitise some of less important material which potentially contains information that is useful for Latvian researchers.

What is your experience with digitised collections of other countries?
I have been trying to find some information about Latvian actor Viļums Vēveris who was living and performing in Cologne (Köln) for two years. I couldn’t find much in the German digitised collections unfortunately. I presumed that maybe there is some information in the publications that are not digitised or that I just didn’t know where to look or how to search properly. Therefore, I believe, it is quite important that librarians are able consult extensively also about foreign collections, or that support for the foreign researchers is readily available.

Do you see benefits in creating collaborative aggregated resources, like Europeana Newspapers?
Yes, they could be very helpful. It could make comparative studies on various subjects much easier.

In your opinion, to what extent digitisation may change the humanities? Are there new topics, new requirements for research?
Naturally, there will be changes. We can work with much more material. It can be seen very clearly now at what point a new term appears in the language and culture, data can be made into diagrams, if needed. The way we see the task of the researcher changes. What is a true researcher? We used to see him as somebody who has rifled through every publication and has spent most of the days of his life in the library, turning dusty pages of newspapers.

Interview by Anda Baklāne, National Library of Latvia, April 2014

]]>
http://www.europeana-newspapers.eu/karlis-verdins/feed/ 0
Q&A with newspapers researchers: Toine Pieters http://www.europeana-newspapers.eu/qa-with-newspapers-researchers-toine-pieters/ http://www.europeana-newspapers.eu/qa-with-newspapers-researchers-toine-pieters/#comments Mon, 31 Mar 2014 07:11:05 +0000 http://www.europeana-newspapers.eu/?p=2670 Continue reading Q&A with newspapers researchers: Toine Pieters]]> Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering? This month we’re interviewing Toine Pieters.

Toine Pieters: ‘Digitisation opens up a world of opportunities, but we must preserve the printed originals for verification’.
Toine Pieters: ‘Digitisation opens up a world of opportunities, but we must preserve the printed originals for verification’.

 

“European libraries must secure free access for researchers”

You can find Professor Toine Pieters on the fertile grounds where the life sciences meet the humanities and social sciences – with a background in pharmacology and a PhD in Social Studies of Science he teaches history of the life sciences at Utrecht University. Pieters: ‘Medicines have always been at the forefront of the public debate – both as focal points of cultural enthusiasm and as the locus of public controversy. The utopia of the wonder drug and the dystopia of poisoned minds or bodies represent the opposite ends of the range of attributed meanings.’

‘One of my current projects is Translantis, research into so-called ”reference cultures”. How do we perceive foreign cultures which are dominant in a certain field or age? The Americans have long dominated the pharmaceutical debate. Before them it was the Germans. It is interesting to study tell-tale signs of reference cultures in our public sphere. And see how the shift occurred from the Germans to the Americans.’

 

Newspapers resonate with public sentiment”

 ‘Newspapers do not only reflect the news of the day, but also the way contemporaries perceived those events, in background articles, editorials and letters to the editor. Plus: they constitute a very large corpus. This makes newspapers a true treasure trove for studying public discourse, the interaction between science, culture and the economy.’

 

Digitisation has opened up fabulous new research opportunities”

‘In the analogue days, any number of researchers could only analyse a fraction of the information that was out there. Also, the research question very much determined the selection. Now software allows us to work with millions of pages. By combining words and expressions, machines uncover patterns that we never even suspected were there.’

‘For example, I researched the duplicitous attitude of the Dutch towards drugs before World War II. At home their religion forbade the use of drugs, but in the Dutch colonies they were actively engaged in the production and trade of cocaine and opium. Somehow the Chinese living in the Netherlands were also involved, but we did not quite know how. By analysing so-called “hidden debates” in the newspapers,  we were able to uncover a pattern revealing that the Chinese were identified as “the other” to whom the Dutch laws did not apply as much. Their involvement in the drug trade as users and as dealers was in line with the economic interests of the Dutch Government and turned out to be the key to the double standard.’

An illustration of the Dutch duplicitous attitude towards drugs as reflected in the Nieuws van den Dag voor Nederlandsch-Indië [News of the Day for the Dutch East Indies], 31 January 1929. ‘Commissioners of the Opium Factory in Naarden, among whom six of the most esteemed and honourable businessmen in the Netherlands, have sent a declaration to the press stating that they have not violated any Dutch laws, because the authorities have not forbidden opium exports to China.’
An illustration of the Dutch duplicitous attitude towards drugs as reflected in the Nieuws van den Dag voor Nederlandsch-Indië [News of the Day for the Dutch East Indies], 31 January 1929.
‘Commissioners of the Opium Factory in Naarden, among whom six of the most esteemed and honourable businessmen in the Netherlands, have sent a declaration to the press stating that they have not violated any Dutch laws, because the authorities have not forbidden opium exports to China.’

 

“We are still in the pre-history of working with big data”

‘But there are pitfalls as well, of course. With big data at our disposal and prototypes of software to work with them, we have yet to develop corresponding methods of source criticism and heuristics. First of all, the available collections are not representative. Libraries have made the selection for us – sometimes consciously, sometimes because they had no choice. Many publishers expect to make money from their digital archives and therefore will not provide free access, even if it is in the public interest. For example, the Dutch post-World War II newspaper collection has a strong bias towards smaller newspapers. All but one of the national papers are missing from the collection. That greatly affects any research outcome.’

 

“OCR quality varies”

‘Another obstacle is the varying quality of the OCR. There are quite a few issues with the earliest scans. Such are the dialectics of lead. Our technical means have greatly improved in recent years, but even present-day OCR is far from perfect. Historical texts continue to confound OCR equipment.’

‘Improving low-quality OCR is a tremendous challenge. Post-digitisation OCR correction tools are being developed, but their usefulness is limited. Sometimes rekeying the content is the only option.  As resources are limited, I think we will need a massive crowd-sourcing effort to get that job done. In the past the KB was not particularly receptive to the idea of crowd sourcing, but hopefully that is changing now. Crowd sourcing is the way to go, it is absolutely necessary.’

 

“Do not throw away the paper originals!”

 The issues we now have with OCR from the past only go to show that it is very, very unwise to do away with paper originals once they are digitised. Governments may think it is a way to save money, but for research it is a downright disaster. In view of the quality issues with OCR we must always be able to go back to the originals to verify whether the digital copy has, in fact, the whole story. For academic research we need the facts, nothing but the facts. ’

 

“Multilingualism at the European level”

 ‘A European aggregation such as Europeana is, of course, a great idea. For the first time in history we have the opportunity to do transnational comparative research. Multilingualism is an obstacle we can overcome, I think, although there is lots of work to be done. I myself am involved in the HERA project which involves newspaper collections from the Dutch KB, the National Library of Luxemburg, some German libraries and the British Library. We have developed a demonstrator for Dutch-German bilingual text mining called BILAND and are now proceeding towards trilingual text mining. Also, I am co-organising a two-day HERA-sponsored seminar on Mining Digital Repositories in a European context in April at the Dutch KB. At the seminar, we will discuss all angles of transnational and multilingual text mining with a select group of researchers [link to on-line report to follow]. Again, this is entirely new ground we are breaking; it is very exciting, but we must develop new methods and rules for such research.’

 

“Libraries should secure free access for researchers at the European level”

‘Looking at the future of newspaper collections and of Europeana, I of course hope that they will keep building and improving the collections. But I see another important task for the libraries of Europe: to negotiate free access to the collections for research purposes.’

‘I think the researchers themselves will develop most of the tools they need, as they know exactly what their research questions are. But securing access must be done at the aggregation level. I understand perfectly well that publishers need to make money by selling access to their archives to the general public. But there must be some way to negotiate unrestricted access for research purposes, in the public interest. There is a crucial role for libraries and Europeana to play here.’

 

Text and photograph Inge Angevaare / March 2014

 

]]>
http://www.europeana-newspapers.eu/qa-with-newspapers-researchers-toine-pieters/feed/ 1
Q&A with newspaper researchers: Wiebke Schulz http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-wiebke-schulz/ http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-wiebke-schulz/#respond Mon, 03 Feb 2014 11:07:44 +0000 http://www.europeana-newspapers.eu/?p=2457 Continue reading Q&A with newspaper researchers: Wiebke Schulz]]> Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering?

To answer that question, we’re interviewing researchers about their work with historic newspapers.

Wiebke Schulz used job advertisements in historic newspapers as one source when researching her PhD thesis on the Careers of men and women in the 19th and 20th centuries.

Wiebke SchulzCan you tell us about your work?
In my dissertation I studied social mobility over the life course in the past and especially the long-term development of social mobility. Sociological theories claim that during the period of modernisation, roughly in the 19th and 20th centuries, tremendous social changes happened and these are assumed to have had an impact on the mobility outcomes of individuals.

In the period before modernisation, people were born into a certain family and social class. There was not much opportunity to improve their social position over the course of their career. Modernization theories argue that, due to the changes that occurred, people had more chances to have an occupation which was different from that of their parents and to make life choices that were different from those of their parents. My goal was to provide an empirical test of these ideas. My dissertation presents the first ever historical description and test of classic theories in the field of status attainment over the working life of the Dutch population at large.

How did historic newspapers fit into your work?
In one chapter of my dissertation, we wanted to focus on the role of the employer. According to modernization theories, employers would increasingly recruit based on qualifications rather than background characteristics such as religious affiliation and social class. We therefore needed a source that covered a variety of occupations, employers and a long period of time. Job advertisements from historic newspapers are an excellent source. Job ads include information on the occupation and tasks that a potential employee has to execute and information on requirements.

An 1894 advertisement for job openings, in the Dutch paper De Telegraaf.
An 1894 advertisement for job openings, in the Dutch paper De Telegraaf.

Did you use paper newspapers or digitised archives, and how easy was the process for you?
We used the digital newspaper archive of the Royal Library of the Netherlands and the archive of the Leeuwarder Courant.

The main challenge for us was related to the types of newspapers that were digitised. We needed papers that covered different social groups so we wanted to include the newspaper read by the socialist labourer but also the newspaper read by the Protestant middle class. It was a challenge to collect a sample of advertisements that allowed us to make general claims on the recruitment of employees in the Netherlands via job advertisements.

What would you like to see for digitised historic newspapers in future?
The historic newspapers we used were digitised but they weren’t processed in a way that allowed us to could search them by keyword. We had to go through every newspaper and re-type the content of the advertisements. That was quite a lot of work. Improvements in this area would be really helpful.

Will you continue to research aspects of social mobility?

Yes, there are so many interesting questions to be studied, in our current society but also in a long term perspective like we did with the newspaper ads. My co-authors  Marco van Leeuwen & Ineke Maas are planning to continue doing research using historic newspapers as a data resource. Their website gives more information about their research projects.

Wiebke Schulz (1983) studied sociology at Bremen University, the University of Liverpool and Utrecht University. Her dissertation “Careers of men and women in the 19th and 20th centuries” was conducted at the Interuniversity Center for Social Science Theory and Methodology (ICS) in Utrecht. Currently, she is employed as post-doctoral reasearcher in the TwinLife project (Bielefeld University/DIW Berlin).

]]>
http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-wiebke-schulz/feed/ 0
Q&A with newspaper researchers: Leon Saltiel http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-leon-saltiel/ http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-leon-saltiel/#comments Tue, 28 Jan 2014 09:30:03 +0000 http://www.europeana-newspapers.eu/?p=2453 Continue reading Q&A with newspaper researchers: Leon Saltiel]]> Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might academics use the material that we’re gathering?

Leon panelTo answer that question, we’re interviewing researchers about their work with historic newspapers. Our most recent conversation was with Leon Saltiel, who is using newspapers held in local and foreign archives to research the period of World War II in Thessaloniki, Greece.

Q. Can you tell us a little bit about your work?

I am researching the period of the Second World War in Thessaloniki, Greece’s second largest city, and I am trying to examine how the city operated during that time. This includes the institutions, the elites, the government and the German occupiers. There is a special focus on the Jewish community. Around 20-25% of the city’s citizens were Jewish and more than 90% were eventually sent to the Auschwitz death camp. That’s my premise: trying to map life in Thessaloniki, identify the main actors and how they were behaving at that time.

Registration of the Jews of Thessaloniki, July 1942, Eleftherias Square. (German Federal Archive / Wikimedia Commons).
Registration of the Jews of Thessaloniki, July 1942, Eleftherias Square. (German Federal Archive / Wikimedia Commons).

Q. Why are newspapers so important as a source of information for you in relation to this research?
Newspapers are crucial to trying to figure out the daily life in the city. You can’t do the research without newspapers. They document each day. For my research, there are two newspapers which were in circulation at the time and thus are contemporary to the events. For me that is key: being contemporary means that there is no hindsight, only the facts as they existed at the time are present. Having said that, both newspapers were controlled by Nazi propaganda at the time. This means that some of the facts are skewed and some information is not there but still you find a lot of material that is very useful for research and often through those articles I am able to find leads to testimonies and other sources.

Q. Are digital versions of these newspapers available or are you using the original print copies?
Both of them are paper newspapers. I have to go to the libraries that hold them, go through the dates that I am interested in, take photos and create a list of press clippings. It takes a lot of time. Unfortunately these newspapers aren’t yet in a digital form. Hopefully soon they will be digitised to facilitate further research on this period. It is important to digitise these newspapers because only two unique copies exist at the library. As time goes by they are being physically destroyed. We need to digitise them to preserve them for future historians. Otherwise in 5-10 years no one will be able to use them and this unique resource is going to be lost. Already some of the issues are in very bad shape.

Q. If you had the choice today to do your research using a digital archive or the original copies, which would you choose?
I would definitely use the digital archive. I have used digital archives for other research projects and they are so easy to use. You can access them from your laptop or your desk, at any location. The text is searchable so you can type in a keyword and find things that you never thought you could find. You also don’t have to worry about damaging the physical newspaper. The newspapers I study from 1943 had been printed on poor paper due to the Second World War and every time you turn a page you can hear the page being ripped. In a digital archive there is no damage caused to the real paper.

Q. In your previous work with digital archives, what techniques have you used?
The digital archives I used did not have the possibility to do text mining, nor did I have that analytical element in mind so I was doing basic keyword searches. This can be problematic because the Greek letters have a lot of accents which were only simplified in the beginning of the 1980s. These aren’t read very well by computers so you can miss things with the keyword search or get results that shouldn’t be there. I also sometimes looked for dates to see what the press was reporting in the time around a certain event.

Q. The research you’ve done is on a local scale. If we look forward, what potential do you see for a pan-European archive such as the one being built by Europeana Newspapers? Could you, for example, extend your thesis?
It’s a good question. My thesis is what we call micro-history. It takes a very specific look at a single city over a period of 1-2 years but there are many inter-related events that happened elsewhere and you do need context in your research so newspapers in other countries could be interesting. Maybe some trials took place after the War that relate to characters I am researching. Also, newspapers sometimes have photographs of different events or people. Finding images can be difficult but newspapers are a great source of photographs.

Leon Saltiel is a PhD candidate at the Department of Balkan, Slavic and Oriental Studies of the University of Macedonia, in Thessaloniki, Greece. He was a Fulbright Scholar at Georgetown University, earning a master’s degree in Foreign Service, and he holds a bachelor’s degree in international relations from the University of Macedonia. Leon has been named a Marshall Memorial Fellow of the German Marshall Fund, a Junior Fellow in Diplomacy of the Institute for the Study of Diplomacy and a Student Honoree of the American Academy of Achievement.

]]>
http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-leon-saltiel/feed/ 1
Q&A with newspaper researchers: Bob Nicholson http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-bob-nicholson/ http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-bob-nicholson/#respond Tue, 21 Jan 2014 11:39:58 +0000 http://www.europeana-newspapers.eu/?p=2435 Continue reading Q&A with newspaper researchers: Bob Nicholson]]> Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering?

Bob Nicholson

To answer that question, we’re interviewing scholars about their work with historic newspapers. Our most recent conversation was with Bob Nicholson, a historian of 19th-century popular culture at Edge Hill University.

He used newspaper archives extensively to research his PhD thesis, which included an examination of how content — and in particular jokes and slang — moved between America and Britain in the 1800s.

Jokes in newspapers

How did you discover the link between your research topic and newspapers?

It was quite accidental and goes back to before we had access to digital newspapers. I was in a local library helping my dad to do research. He’s also a historian. I was only about 17 years old and we were looking through microfilms related to his research when, completely by accident, I came across a column of imported American jokes. I never expected to see that in a newspaper from the 1880s. I thought: “This is unusual. Where have these come from?”

Over the next few years, bit by bit, I started to map these connections and what I discovered is that there is an enormous amount of content from America in British newspapers and vice versa. It was an accidental discovery but what digital archives allowed me to do was to really turn that accidental discovery into something much bigger. I wrote an essay about it for my MA. That turned into my PhD and I’m now turning that into a book.

You clearly found plenty of material to work with!

Yes, and this is perhaps the problem with digital archives more generally. There is just so much stuff. Every time I find a connection it leads me on to a new research topic. In the end I have enormous amounts of research and I never get anything finished because I always find something new. It’s a nice problem to have in the sense that I’m never going to run out of ideas but it also means that turning it into a book or a cohesive piece of research is tricky.

I remember that one of the first things I did at the start of my PhD research was to put the word America into one of the digital newspaper databases. It inevitably came back with millions and millions of hits and I could just see the next 10-20 years of my life standing there before me and thinking “I’m never going to get through all of these”.

How has this work affected your view of newspapers?

There were two things that changed for me. The first was my own approach to newspapers. When I started my own idea was to do what a lot of historians do, to mine them for information. I wasn’t really interested in the work that newspapers perform themselves. I just thought “wow, this is an enormous bank of 19th century culture that I can look through to find what the Victorians thought about America”. After a few months, I realised the agency of newspapers and the power they had to shape this process. I realised just how powerful newspapers were and the influence they had on 19th century culture.

The content also surprised me. Almost every time I go into a newspaper I find something unexpected and that’s one of the great pleasures of using them. They contain an extraordinary amount of unusual things. All of the world is in there. Every aspect of 19th century culture, just about, is captured in some form either accidently or deliberately. It means that they are an almost endlessly useful resource. As a cultural historian, I think I’ll be using them for the rest of my life.

Terrified To Death By A Donkey: Just one example of the unusual stories that Bob Nicholson has found in historic newspapers. This appeared in the Illustrated Police News, 1883.
Terrified To Death By A Donkey: Just one example of the unusual stories that Bob Nicholson has found in historic newspapers. This appeared in the Illustrated Police News, 1883.

How would you compare them to other sources of information such as books and journals? Are there certain aspects of newspapers that just can’t be replicated anywhere else?

Newspapers don’t exist in a sealed off world. They’re on a shelf right next to a magazine or a book and people walk past talking about them so I wouldn’t look at newspapers just on their own but I do think they give us something quite distinctive.

One of the things is the fact that they are appearing every day or every week. As historians we’re interested in change and there is nothing that is more useful than being able to take a text and see how, on a daily or weekly basis, things change within it. A book is a single text in a moment in time whereas newspapers allow you to track change in a number of ways.

The phrase that newspaper historian Martin Conboy uses for newspapers is that they have a heteroglossia, a number of voices. I love that, the fact that they’ve been written by a number of people in a number of styles. This makes them quite unique.

I’d like to go on to how you actually do your work. You primarily use digital archives. Is that correct?

It is, yes. My entire research is dependent on them now. I rarely use paper or microfilm copies unless I really have to. Since I started my MA around the same time as the British Newspaper Archive was launched, I could design my doctorate research entirely around the use of digital archives, with a view to exploring the methodological possibilities. Of course I did look at some paper copies (there are some absolutely vital newspapers that I could only get on microfilm) but for the most part I’ve designed everything I’ve done since 2007 around digital archives.

What is your work process? Do you perform simple searches or do you download material and mine it for information?

A typical joke column in a British newspaper, from 1889.Most British newspapers are behind paywalls, which means that as researchers we don’t have the ability to download material and mine it. I’ve had to use the existing interfaces. Thankfully I was able to do reasonably good keyword searching within the archives.

I also developed a very basic sort of text mining myself where I did searches and tracked the frequency of certain terms appearing within the newspapers. It was very laborious. In the end it took me a day to perform each search. It would have taken me about 30 seconds using a proper tool but I did find some very useful material for illustrating big changes relative to my PhD. This work was critical because, as I said before, if you have a million newspapers mentioning the term “America” there’s no way you can read them all so I needed a way to “distant read” them and uncover the patterns. I would love to have better tools to do that automatically.

Do you ever worry that you may have missed material because you were limited to what had been pre-selected to appear in digital archives?

There were some papers that I knew I wanted to look at which weren’t digitised so I did what we all used to do and spent a week in the archives looking at material. I bought things online if I had to. But the fact is that I could access somewhere in the region of 1,000 titles from the 19th century from home. There’s no way that even five years earlier I would have been able to look at 5% of those things for my PhD so I wasn’t too anxious. I knew that the benefits of digital archives massively outweighed the problems.

If you look forward 10 years, what would you like to see?

What I would really love is to liberate newspapers from behind the paywalls so that we could start doing more interesting things with them. The best research projects that I’ve encountered in the digital humanities are completely open and allow us to do incredible things. I’d like to be able to do more proper text mining, tracking the frequency of words automatically. That’s at the top of my list currently. For me, we have enough content. It’s more about developing the tools to make the most of it.

I’d also like to ask about extending your work across borders. The digital archive that we’re building is pan-European. What potential do you see for that body of material?

The work I’m doing is very trans-national so I’ve had to make links between different databases in ways that are sometimes quite tricky. I see enormous value in an archive that breaks down national boundaries automatically, where I can search for content from a range of countries. There’s huge potential to look at newspapers on a pan-European perspective.

The problem for me and I imagine a lot of British and American scholars will be the language barrier. My language skills are terrible and I don’t know whether I’ll actually be able to make the best use of those resources but I can see enormous potential. If we don’t have to learn to read the languages at a native level, we might still be able to search for certain words and then through translation software we might still be able to find those connections.

Maybe you could also collaborate with researchers in other countries and overcome the language barrier in that way.

That’s obviously what digital archives allow us to do. They break down systems where earlier you’d have to sit, for example, in a specific room in London to do research. That change should, at least in theory, allow us to have these sorts of collaborations. One of the weird things about historians is that we don’t seem to work in those collaborative lab-based environments as much as in other disciplines but if I could find somebody who’d be willing to spend their life looking at 19th century jokes in German that would be great.

Before we close this interview, I just have to ask. What’s the best joke you’ve found during your research?

I always get this question! The problem is that when you’ve read about 10,000 jokes you lose any kind of ability to decide what is funny anymore. It warps your sense of humour. Also, jokes in newspapers aren’t really designed to be told orally. I do have one joke that I’ve told at tons of conferences and almost invariably no one laughs at it. It creates this incredibly awkward silent moment but I can share it with you.

CHICAGO WOMAN: What do you charge for securing a divorce?

CHICAGO LAWYER: $100 ma’am, or six for $500!

Bob Nicholson can be found online via his website The Digital Victorianist. You can also follow him on Twitter: @DigiVictorian.

]]>
http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-bob-nicholson/feed/ 0
Q&A with newspaper researchers: Matthew Rubery http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-matthew-rubery/ http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-matthew-rubery/#respond Wed, 08 Jan 2014 07:30:22 +0000 http://www.europeana-newspapers.eu/?p=2407 Continue reading Q&A with newspaper researchers: Matthew Rubery]]> Within the Europeana Newspapers project, we often speak of the value of historic newspapers for the academic community but how exactly might a researcher use the material that we’re gathering?Matthew Rubery photo

To answer that question, we’re interviewing researchers about their work with historic newspapers.

Matthew Rubery was the first to respond to our questions. He teaches at Queen Mary University of London and is the author of The Novelty Of Newspapers: Victorian Fiction after the Invention of the News. This book explains how newspapers shaped the development of Victorian fiction.

Your book highlights the importance of newspapers as an information source. How did you discover the link between your research topic and the information recorded by the press?

Rubery Book CoverMy first encounter with the 19th-century press was a memorable one. While browsing the archives, I was surprised to discover that the entire front page of the London Times was covered with tiny advertisements rather than news. Where were the attention-grabbing headlines that sought, in one editor’s words, to strike the reader between the eyes? Clearly the press has come a long way since those days. My book traces the impact such dramatic changes to the press had on the work of Charles Dickens and a whole bunch of other novelists who mined journalism for fictional inspiration.

Was this the first time that you had used newspapers as a source of information for research, and did it change the way that you view newspapers?

Anyone who works with old newspapers knows how easy it is to get lost in them. They’re full of human interest stories that continue to fascinate us centuries later. I began working with newspapers to gain historical context but quickly found myself absorbed by the narratives unfolded in their pages: tales of scandalous affairs, shocking crimes, disasters at sea, and more. Journalists are very skilled at converting intelligence into readable narrative forms so that even something as simple as an obituary contains the story of an entire life.

For centuries newspapers have informed as well as entertained. We should remember that many of Europe’s greatest novelists wrote for the press. Classic novels such as Dickens’ Oliver Twist, George Eliot’s Romola, and Arthur Conan Doyle’s The Hound of the Baskervilles first reached audiences in Britain through the pages of newspapers and magazines.

This was true in other parts of Europe too. Eugene Sue’s Les Mystères de Paris, one of the most influential novels of the nineteenth century, was published in 90 installments in a French newspaper. Newspapers are commonly thought of as disposable but it’s hard for me to think of them that way when they’re full of such great stories.

What specific types of information did the newspapers contain that you found valuable, and why were these important for you?

The Novelty of Newspapers examines the interplay of journalism and the novel. Its chapters are organised according to types of newspaper features: shipping news, agony columns, leading articles, personal interviews, and foreign correspondence. I chose these categories in order to show how public media enabled the expression of private feeling. A few lines about a ship lost at sea, for instance, could be incredibly emotional when read by someone who knew, or loved, one of that ship’s passengers. To take a dramatic example from fiction, Esther Summerson is scarcely able to read the account of Dr Woodcourt’s shipwreck through her tears in Dickens’ Bleak House. Newspapers were a great way of bringing Europeans into contact with the rest of the world.

In terms of your work process, what kinds of newspapers did you use and how did you use them? 

My book should have a disclaimer that it was written before the arrival of digital collections. A book that took me years to write could now be completed in weeks! It’s easy to forget how difficult it used to be to access newspapers.

Most were thought to be ephemeral. The saying comes to mind: “Yesterday’s news is tomorrow’s fish and chip wrapper.” Few libraries bothered to preserve them; those that did generally had cumbersome microfilm copies.

While writing my book, I spent a year in London working at the British Library’s newspaper repository in order to gain access to material that couldn’t be found in the United States. In fact, newspapers are largely responsible for the fact that I live and work in Britain today. I warn my doctoral students to be careful where they do archival research since they may end up spending the rest of their lives there!

How was your research affected by the format of newspapers that you used, in both a positive and a negative sense?

We’re witnessing dramatic changes in access to information through digital archives. Newspapers are so accessible nowadays that my students use them in the classroom (instead of the flimsy, smudged sheets that I used to pass around).

The difference lies not just in access, of course, but in the conversion of a massive amount of print into a searchable resource. We can search millions instead of hundreds of pages. This holds the potential to make connections across newspapers in ways previously unimaginable to the lone scholar thumbing through page after page of the broadsheets.

The boom in digital resources means that we have gone from too little to too much information almost overnight. This makes it all the more important that we avoid “cherry picking” information without grasping the conventions governing the press at different moments in history. In my case, I would never have made the connections I did without viewing the physical format of the newspaper and reflecting on the way people actually read it.

Scanning the newspaper’s front page, for example, brought me into contact with the mysterious world of personal advertisements. One of my favorites is from 1850: “THE ONE-WINGED DOVE must DIE unless the CRANE RETURNS to be a shield against her enemies.” What was to be the abandoned dove’s fate? This story would have been lost to me were it not for the serendipity of browsing.

Looking forward, can you imagine other topics or areas of research that might be stimulated by the digital archive that we are building and by digital newspapers in general?

One of the most exciting—and controversial—trends in the humanities right now is the shift from “close reading” of a single text to “distant reading” of massive amounts of information. This approach only became possible after the conversion of millions of books, magazines, and newspapers into a machine-readable resource through digital repositories such as Google Books.

When done responsibly, this approach should help us to map the evolution of ideas, language, and culture on an unprecedented scale. For instance, we might be able to chart formal innovations in the press such as the emergence of headlines on the newspaper’s front page. Data mining makes perfect sense for an industry that has always prided itself on quantity as much as quality.

Matthew Rubery is the author of The Novelty Of Newspapers: Victorian Fiction after the Invention of the News. His book is currently available as a hardcover from Oxford University Press and will shortly also be available as a paperback.

]]>
http://www.europeana-newspapers.eu/qa-with-newspaper-researchers-matthew-rubery/feed/ 0