Posts

Showing posts with the label Big Data

Using Big Data to Solve Social Science Problems

Image
Curtis Jessop  is a Senior Researcher at NatCen Social Research and is the Network Lead for the NSMNSS network On Wednesday 29 th June I attended a roundtable hosted by our network partners SAGE on using big data to solve social science problems. It was a great day, with contributions from leading researchers and lots of discussion of some of the key issues of working with big data in social science. Jane Elliott began with an overview of the ESRC’s Big Data Network. She identified the difficulties with data access that earlier phases had faced, but also highlighted key challenges that big data social science currently faces: 1. Methodological Can we apply the same qualitative techniques/statistical inferences we have in the past? Are social scientists (falling) behind in using machine learning & algorithms? What are the implications of these methods? 2. Relevance of research Making sure we use big data to answer pertinent social science questions, and not just focus on meth...

Making the most out of big data: computer mediated methods

Patrick Readshaw is a Media and Cultural Studies Doctoral Candidate at Canterbury Christ Church University. Patrick is interested in social media as an alternative and empowering source of information on current events, free from the constraints of other agenda-setting media forms. You can contact Patrick by email on p.j.readshaw68@canterbury.ac.uk     When I was asked to write a blog for NSMNSS, I was certainly excited and being my first post of this kind I was suitably anxious about the prospect. However, my ongoing thesis has never ceased to provide interesting discussions with individuals in linked or parallel fields relating to social media. The main caveat in these discussions is that I often have to try not to over complicate things. With that in mind and my ham-fisted introduction out of the way I want to take some time to break down the value of so called “new media systems” like Twitter and the how I personally go about dealing with the data I collect.  Sinc...

The changing nature of who produces and owns data: How will it impact survey research?

Image
Brian Head is a research methodologist at RTI International. This post first appeared on SurveyPost on 20 May, 2014. You can follow Brian on Twitter @BrianFHead . Survey researchers have become interested in big data because it offers potential solutions to problems we’re experiencing with traditional methods. Much of the focus so far has been on social media (e.g., Tweets), but sensors (wearable tech) and the internet of things (IoT) are producing an increasingly rich, complex, and massive source of data. These new data sources could lead to an important change in how individuals see the data collected about them, and thus have ramifications for those interested in gathering and analyzing those data. Who compiles data? Quantitative data about people have been gathered for millennia. But with technological advances and identification of new purposes for it, the past 100 years have seen significant increases in the amount of data produced and collected—e.g., data on consume...

You Are What You Tweet: An Exploration of Tweets as an Auxiliary Data Source

Image
Ashley Richards is a survey methodologist at RTI International. This post first appeared on SurveyPost on 29, July 2014.   Last fall at MAPOR  , Joe Murphy presented the findings of a fun study he did with our colleague, Justin Landwehr, and me. We asked survey respondents if we could look at their recent Tweets and combine them with their survey data. We took a subset of those respondents and masked their responses on six categorical variables. We then had three human coders and a machine algorithm try to predict the masked responses by reviewing the respondents’ Tweets and guessing how they would have responded on the survey. The coders looked for any clues in the Tweets, while the algorithm used a subset of Tweets and survey responses to find patterns in the way words were used. We found that both the humans and machine were better than random in predicting values of most of the variables. We recently took this research a step further and compared the accuracy of ...

Save the dates! Upcoming tweet chats

There have been lots of interesting discussions and topics floating around about new social media in the social scienes. What better way to share than to host some tweet chats! See below for the dates and the topic we will cover for each tweet chat. All times are London time. Tuesday 7th October, 2014 at 5pm: Representativeness of online samples Including: What are the geographical inequalities in contributions across different social media platforms? What approaches can we take to address this? How can we weight twitter data? How can we learn about demographics of people on social media, such as age, gender, employment? Monday 17th November at 5pm: Ready, set, research!: accessing funds and data Including: You have an idea for a study, how do you go about funding it? What funding streams are available? What are the regulations/restrictions of accessing different streams? How do we get our hands on big data sets from the likes of Google and Twitter? Tuesday 9th December at 5pm: The ch...

Understanding Geert Lovink’s book “Networks without a cause: A critique of social media”: A video by Akin Olaniyan

Akin Olaniyan is a student in the MA in Social Media at the University of Westminster Watch the video interview here!  https://www.youtube.com/watch?v=GV-PBy3iaDk He sounds very much like a techno-pessimist. When I first read Professor Geert Lovink’s book, ‘Networks without a cause: A critique of social media,’ the first thing that struck me was the sense of despair that runs through all the chapters. With chapters devoted to Facebook and the crisis of identity, big data and the ‘Googlisation of our lives,’ and a proposition to divorce the study of social media from media studies, there appeared to be no other way to understand Geert Lovink. And that was the starting point of my interview with him conducted via Skype. As you can see in the video ( https://www.youtube.com/watch?v=GV-PBy3iaDk ), when I asked him what was the source of the despair I felt running through his book, Geert describes his frustration that the Internet has become too centralized and cites cloud computing a...

Keeping up with technology: What is “scientific lag” and can we proactively reduce it?

Image
In 2011 then Census Director Robert Groves  wrote  on the Census Director’s Blog about the burgeoning volume of “organic data”—data that, as opposed to “designed data,” have no meaning until they are used (surveys are a primary example of the latter). He noted that finding ways to combine these two types of data to increase the “information-to-data ratio” was a challenge, but also represented the future of surveys. Using terms identified as “big data descriptors” in Groves’ piece, as well as a few other terms I think qualify, I put together the graph below to show the number of AAPOR presentation titles between 2010 and 2013 that contain a big data descriptor. 1,2 One take away is the increased interest researchers have shown in big data over the past few years. An equally important lesson is that almost all of the attention big data has received from AAPOR members—at least measured by the number of presentations they’ve done—has been on social networking sites (SNS). I f...

Using “Small Data” to Improve the Use of “Big Data”

Image
This post was first published on Survey Post on Feb. 3rd, 2014. Recently, I attended two statistical events in the Washington, DC, area: one was the 23rd   Morris Hansen Lecture   on “Envisioning the 2030 U.S. Census”; the other was the   SAMSI workshop   on “Computational Methods for Censuses and Surveys.” “Big data” was a popular keyword at both events and stirred up discussions on how to utilize it (such as from administrative records and online data sources) for current government statistics, especially when combining big data with  traditional survey data. Statisticians are exploring new ways in which big data can be used. The US Census has initiated investigations on  using administrative records in the 2020 Census . The National Center for Health Statistics (NCHS) has identified some  research opportunities combining multiple data sources . University-based researchers  have launched studies on the use of Google trends and other online da...