Source Repository: Other Borealis Collections / Publication date: / Author: Ruest, Nick - Lunaris - Discover Canadian Research Data Search Results

#healthcanada #NACI #fordnation #medicalfreedom #covid19 #covid19vaccines #protectourfamilies #protectyourchildren #holdtheline tweets

Borealis

Ruest, Nick — 2022-01-10

Avery Library Historic Preservation and Urban Planning web archive collection derivatives

Borealis

Ruest, Nick; Sala, Christine; Thurman, Alex — 2020-02-16 Web archive derivatives of the <a href="https://archive-it.org/collections/1757">Avery Library Historic Preservation and Urban Planning</a> collection from <a href="https://archive-it.org/home/Columbia">Columbia University Libraries</a>. The derivatives were created with the <a href="https://github.com/archivesunleashed/aut/">Archives Unleashed Toolkit</a> and <a href="https://cloud.archivesunleashed.org/">Archives Unleashed Cloud</a>. The cul-1757-parquet.tar.gz derivatives are in the <a href="https://parquet.apache.org/">Apache Parquet format</a>, which is a <a href="http://en.wikipedia.org/wiki/Column-oriented_DBMS">columnar storage</a> format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See <a href="https://github.com/archivesunleashed/notebooks/blob/master/datathon-nyc/parquet_pandas_stonewall.ipynb">this</a> notebook for examples. Domains <pre> <code class="language-java">.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)</code></pre> Produces a DataFrame with the following columns: <ul> <li>domain</li> <li>count</li> </ul> Web Pages <pre> <code class="language-java">.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))</code></pre> Produces a DataFrame with the following columns: <ul> <li>crawl_date</li> <li>url</li> <li>mime_type_web_server</li> <li>mime_type_tika</li> <li>content</li> </ul> Web Graph <pre> <code class="language-java">.webgraph()</code></pre> Produces a DataFrame with the following columns: <ul> <li>crawl_date</li> <li>src</li> <li>dest</li> <li>anchor</li> </ul> Image Links <pre> <code class="language-java">.imageLinks()</code></pre> Produces a DataFrame with the following columns: <ul> <li>src</li> <li>image_url  </li> </ul> The cul-1757-auk.tar.gz derivatives are the <a href="https://cloud.archivesunleashed.org/derivatives">standard set of web archive derivatives</a> produced by the Archives Unleashed Cloud. <ul> <li>Gephi file, which can be loaded into <a href="https://gephi.org/">Gephi</a>. It will have basic characteristics already computed and a basic layout.</li> <li>Raw Network file, which can also be loaded into <a href="https://gephi.org/">Gephi</a>. You will have to use that network program to lay it out yourself.</li> <li>Full text file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.</li> <li>Domains count file. A text file containing the frequency count of domains captured within your web archive.</li> </ul> Due to file size restrictions in Scholars Portal Dataverse, each of the derivative files needed to be split into 1G parts. These parts can be joined back together with <code>cat</code>. For example: <pre><code class="bash">cat cul-1757-parquet.tar.gz.part* > cul-1757-parquet.tar.gz</code></pre>

Network Data for the Web Archives for Longitudinal Knowledge (WALK) Project

Borealis

Milligan, Ian; Ruest, Nick; Deschamps, Ryan — 2016-08-23

#MakeDonaldDrumpfAgain tweets

Borealis

Ruest, Nick — 2016-03-03

#elxn43 tweets (43rd Canadian Federal Election)

Borealis

Ruest, Nick — 2019-11-23

Tweet ids for final Tragically Hip concert

Borealis

Ruest, Nick — 2016-12-31

#YMMfire tweets

Borealis

Ruest, Nick — 2016-08-21

#thechalkening tweets

Borealis

Ruest, Nick — 2016-04-13

#panamapapers tweets

Borealis

Ruest, Nick — 2016-04-13

#NDP2016 tweets

Borealis

Ruest, Nick — 2016-04-27

#WomensMarch tweets January 12-28, 2017

Borealis

Ruest, Nick — 2017-01-29

#climatemarch tweets April 19-May 3, 2017

Borealis

Ruest, Nick — 2017-05-03

Tweets to Donald Trump (@realDonaldTrump)

Borealis

Ruest, Nick — 2017-12-10 362,464,578 tweet ids for tweets directed at Donald Trump (@realDonaldTrump), collected with <a href="http://www.docnow.io/" target="_blank">Documenting the Now's</a> twarc. Tweets can be “<a href="https://medium.com/on-archivy/on-forgetting-e01a2b95272" target="_blank">rehydrated</a>” with Documenting the Now’s <a href="https://github.com/DocNow/twarc" target="_blank">twarc</a>, or <a href="https://github.com/DocNow/hydrator" target="_blank">Hydrator.</a> <code>twarc hydrate to_realdonaldtrump_20210120_ids.txt > to_realdonaldtrump_20210120.jsonl</code>. Collection notes: <ul> <li>Tweets from May 7, 2017 - October 16, 2018 of the dataset used a combination of the Filter (Streaming) API and Search API.</li> <li>The Filter API failed on June 21, 2017.</li> <li>From June 23, 2017 forward only the Search API was used to collect.</li> <li>Collection was done every 5 days on a cron job, and periodically deduplicated.</li> <li>There is a data gap from <code>Tue Jul 28 13:53:50 +0000 2020</code> through <code>Thu Aug 06 09:36:23 +0000 2020</code> due to a collection error.</li> </ul> This dataset also includes a number of derivative csv files from the original <code>jsonl</code> collected. This includes: <ul> <li>A user csv file created with <a href="https://stedolan.github.io/jq/" target=_blank">jq</a> (see below).</li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-user-information" target=_blank">twut userInfo</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-tweet-language" target=_blank">twut language</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-tweet-times" target=_blank">twut times</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-tweet-sources" target=_blank">twut sources</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-hashtags">twut hashtags</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-urls" target=_blank">twut urls</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-animated-gif-urls" target=_blank">twut animatedGifUrls</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-image-urls" target=_blank">twut imageUrls</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-media-urls" target=_blank">twut mediaUrls</a></li> <li><a href="https://github.com/archivesunleashed/twut/blob/main/docs/usage.md#extract-video-urls" target=_blank">twut videoUrls</a></li> </ul> User csv: <code>jq -r '[.id_str, .created_at, .user.screen_name, .retweeted_status != null] | @csv' to_realdonaldtrump_20190130.jsonl > to_realdonaldtrump_20190130_users.jsonl </code>

#elxn44 tweets (44th Canadian Federal Election)

Borealis

Ruest, Nick — 2021-11-08

Wet'suwet'en tweet ids

Borealis

Ruest, Nick — 2020-04-19

Tyendinaga tweet ids

Borealis

Ruest, Nick — 2020-04-19

Derivative data for the Canadian Political Parties and Interest Groups collection

Borealis

Milligan, Ian; Ruest, Nick; Lin, Jimmy — 2015-12-01

#jcdl2016 tweets

Borealis

Ruest, Nick — 2016-06-27

#paris #Bataclan #parisattacks #porteouverte tweets

Borealis

Ruest, Nick — 2015-12-12

#elxn42 tweets (42nd Canadian Federal Election)

Borealis

Ruest, Nick; Library and Archives Canada — 2015-12-07

Limit by map area

Map search instructions

1.Turn on the map filter by clicking the “Limit by map area” toggle.

2.Move the map to display your area of interest. Holding the shift key and clicking to draw a box allows for zooming in on a specific area. Search results change as the map moves.

3.Access a record by clicking on an item in the search results or by clicking on a location pin and the linked record title.

Note: Clusters are intended to provide a visual preview of data location. Because there is a maximum of 50 records displayed on the map, they may not be a completely accurate reflection of the total number of search results.

Search

Search Results

Map search instructions

Search Details

Limit by:

Source Repository Clear

Publication date

Author(s) Clear