Page Not Found
Page not found. Your pixels are in another canvas. Read more
A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.
Page not found. Your pixels are in another canvas. Read more
This is a page not in th emain menu Read more
Published:
This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool. Read more
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool. Read more
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool. Read more
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool. Read more
Short description of portfolio item number 1
Read more
Short description of portfolio item number 2
Read more
“Evaluating Social Media Reach via Mainstream Media Discourse,” Workshop/Meeting, ODU CS PhD Gathering.
“Xenophobic Events vs. Refugee Population–Using GDELT to Identify Countries with Disproportionate Coverage,” Conference (Full/Poster), SBP BRiMS ‘23.
“Less than 4% of Archived Instagram Account Pages for the Disinformation Dozen are Replayable,” Conference (Short), JCDL ‘23.
“Exploring Xenophobic Events through GDELT Data Analysis,” Student Conference (Full), MSVSCC ‘23.
“Supporting Account-based Queries for Archived Instagram Posts,” MS thesis defense, ODU CS.
My master’s thesis titled, “Supporting Account-based Queries for Archived Instagram Posts”. We addressed a fundamental challenge in social media preservation: valuable Instagram content exists in web archives, but users often cannot discover it because they do not know the original post URLs. I developed two approaches to enable users to retrieve archived Instagram posts associated with a specific account, one leveraging WARC revisit records within the Internet Archive and the other using an external index mapping Instagram accounts to their posts. I implemented and evaluated both approaches, highlighting their practical advantages and limitations for web archives.
We developed a conditional random field (CRF) model that combines text-based and visual features to automatically extract metadata from scanned Electronic Theses and Dissertations (ETDs). We evaluated our method using 500 human-validated ETD cover pages. The model significantly outperformed text-only and heuristic approaches, achieving 81.3%–96% F1 scores across seven metadata fields.
We examined the challenges of preserving Twitter content in web archives following Twitter’s June 2020 UI change. We found that many archives could not properly capture the new interface, leading to missing or incomplete content and potentially inaccurate historical records. Using the personal Twitter account of the 45th U.S. President, @realDonaldTrump, we identified missing evidence of misinformation labels and temporal inconsistencies in archived pages. The goal of this study is to caution the researchers to critically evaluate archived social media data before drawing historical conclusions.
This project developed storytelling solutions to improve the discovery and understanding of web archive collections. I contributed to several components of the DSA project, including AIU, Raintale, and MementoEmbed. AIU is a Python library that gathers seed URLs and metadata from web archive collections such as Archive-It, Pandora, and Trove using APIs and screen scraping. MementoEmbed generates archive-aware cards for individual mementos, while Raintale organizes these cards into engaging, shareable stories. Taken together, these tools help transform large and complex web archive collections into accessible and meaningful experiences for archivists, researchers, and the public.
This was a large interdisciplinary research project which was funded by the U.S. Department of Defense Minerva Research Initiative investigating methods for studying hard-to-reach environments (Khayelitsha Site-C, South Africa and Villa Caracas, Colombia). The project brought together an international team across the U.S., Canada, Norway, and Colombia. The work examines methodological and epistemological gaps when only certain research approaches are feasible. Our component focuses on analyzing these sites using public data sources such as global news and social media, alongside complementary methods including surveys, visual sociology, and citizen science across collaborating teams. The project was later discontinued before completion.
This project is titled “Data Science for Social Good: Mining and Visualizing Worldwide News to Monitor Xenophobic Violence” and was funded by the 2022-2023 ODU Data Science Seed Funding Program. We examined xenophobic events related to refugees and migration using the GDELT 2.0 database and APIs. Our research used visualizations to explore patterns in media coverage through two case studies. We also discussed the data analysis process and the challenges of working with GDELT data and its tools.
We developed MetaEnhance, an AI-based framework for automatically detecting, correcting, and standardizing metadata errors in Electronic Theses and Dissertations (ETDs). Using a benchmark of 500 ETDs, MetaEnhance achieved near-perfect F1-scores for error detection and correctness F1-scores ranging from 0.85 to 1.00 for five of seven fields.
We analyzed human and robot access patterns in web archives using Internet Archive’s Wayback Machine and Portuguese Web Archive access logs from 2012 and 2019. We classified user sessions as human or robot based on browsing behavior and then analyzed them to identify navigation patterns and temporal preferences. The study found that robots dominated archive traffic, accounting for up to 98% of requests, while their access patterns became more varied over time. Both humans and robots showed a strong preference for recently archived web pages.
Archived web pages can generate repeated, invisible HTTP requests, creating unnecessary traffic and potentially overwhelming web archive servers. Pages requiring frequent updates, such as sports scores or playlists, were especially likely to cause this problem, particularly when missing resources produce repeated 404 errors. We proposed to use Cache-Control headers to cache 404 responses, preventing unnecessary requests from reaching the archive server. Our results showed that this approach can reduce wasted network and computational resources during archival replay.
This study examines how social media content extends beyond its platform and is integrated into television news. Existing research largely measures impact through intra-platform engagement, overlooking television’s role as a widely consumed and more trusted medium with a distinct audience. As a result, additional reach and influence of social media content beyond its native platform remain significantly underestimated. To address this gap, this work proposes a framework for detecting and analyzing social media references in TV news broadcasts.
Published:
This is a description of your talk, which is a markdown files that can be all markdown-ified like any other post. Yay markdown!
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.