Data Mining and Distant Reading- Charlotte
Data mining is a process that shows patterns within data that would be harder for humans to access or inaccessible without the help of a computer. It helps capture data that can then be reused and repurposed for research. It is a great resource that shows different perspectives of otherwise overlooked and especially marginalized communities and their cultures. But data mining comes with limitations, because it relies on preexisting information it may lack representation and miss information. Too much data mining may even change the results that researchers are looking for
Data mining is a type of distant reading that can be used to analyze text and literature’s metadata. Through distant reading, researchers can study literature using computers, data, and statistical analysis rather than looking at individual books. The process can be described as studying literature as a large collection of data. It is done because Moretti believed that computers can oftentimes catch things that humans miss.
But distant reading cannot be done without close reading. It is true that computers find the facts and data but by doing so, it removes human thoughts and features and can often misinterpret or completely lose the actual contents of a story. It can also leave out the complexity and meaning within the writing of key literature. Ironically, quantitative analysis can’t replace close reading because humans still need to read the literature in order to insert the writing into data.
Information visualization also plays a large role in analysis because graphics can reveal patterns and make them easily accessible for audiences. But it is important to note that graphics must be appropriately designed to the research given, in order to prevent misinterpretation.
Six Degrees of Francis Bacon: Visualized interactive and informational graphic that is a compilation of data about Francis Bacon and the relationship/ connection between him and different people. This is an example of distant reading in a digital humanities project.
Yesterday, Today, Tomorrow: Data visualization and distant reading of a large collection of tweets used to analyze how society emotionally experienced the COVID-19 pandemic.
Charlotte, I like your point about how distant reading can help us notice things that would be difficult to see on our own, while still having some pretty important limitations. I definitely do agree that close reading is still necessary. Even if a computer can tell us that certain words, ideas, or connections appear frequently, we still have to look at the actual material to understand why that pattern matters. I felt this when using Voyant because the visualizations gave me a different way of looking at the text, but they didn’t really explain anything on their own. I also think your point about missing information is important. A large amount of data can seem complete or objective when there may actually be people and perspectives missing from it. Yesterday, Today, Tomorrow made me think about this too, since it can show broad emotional patterns during COVID through tweets, but those tweets still can’t represent everyone’s experience.
ReplyDeleteYes, they really work together ideally!
ReplyDelete