awesome-website-change-monitoring by edgi-govdata-archiving

A curated list of awesome tools for website diffing and change monitoring.

created at May 24, 2017, 5:33 a.m.

Unknown languages

31 +0

482 +0

31 +0

GitHub
node-warc by N0taN3rd

Parse And Create Web ARChive (WARC) files with node.js

created at May 21, 2017, 6 a.m.

JavaScript

9 +0

92 +0

20 +0

GitHub
node-cdxj by N0taN3rd

Parse CDXJ(https://github.com/oduwsdl/ORS/wiki/CDXJ) files with node.js

created at May 18, 2017, 4:45 a.m.

JavaScript

3 +0

0 +0

1 +0

GitHub
flameshot by flameshot-org

Powerful yet simple to use screenshot software :desktop_computer: :camera_flash:

created at May 10, 2017, 7:44 p.m.

C++

205 +0

23,358 +42

1,507 +3

GitHub
ArchiveBox by ArchiveBox

🗃 Open source self-hosted web archiving. Takes URLs/browser history/bookmarks/Pocket/Pinboard/etc., saves HTML, JS, PDFs, media, and more...

created at May 5, 2017, 8:50 a.m.

Python

171 +1

20,012 +65

1,089 +3

GitHub
wasapi-downloader by sul-dlss

Java application to download WARCs from WASAPI

created at April 28, 2017, 9:15 p.m.

Java

22 +0

6 +0

4 +0

GitHub
tikalinkextract by httpreserve

Tika based link (URL) extractor for httpreserve

created at April 3, 2017, 2:35 a.m.

HTML

4 +0

8 +0

1 +0

GitHub
har2warc by webrecorder

Convert HTTP Archive (HAR) -> Web Archive (WARC) format

created at March 16, 2017, 12:14 a.m.

Python

7 +0

42 +0

3 +0

GitHub
warcio by webrecorder

Streaming WARC/ARC library for fast web archive IO

created at March 6, 2017, 6:17 p.m.

Python

22 +0

349 +0

55 +1

GitHub
monolith by Y2Z

⬛️ CLI tool for saving complete web pages as a single HTML file

created at Feb. 20, 2017, 7:47 a.m.

Rust

62 +0

10,140 +68

287 +2

GitHub
fbarc by justinlittman

A commandline tool and Python library for archiving data from Facebook using the Graph API.

created at Feb. 14, 2017, 11:45 p.m.

Python

16 +0

77 +0

11 +0

GitHub
WarcPartitioner by helgeho

Partition (W)ARC Files by MIME Type and Year

created at Feb. 13, 2017, 3:45 p.m.

Java

2 +0

1 +0

1 +0

GitHub
archivenow by oduwsdl

A Tool To Push Web Resources Into Web Archives

created at Feb. 9, 2017, 12:29 p.m.

Python

21 +0

392 +0

41 +0

GitHub
solrwayback by netarchivesuite

A search interface and wayback machine for the UKWA Solr based warc-indexer framework.

created at Feb. 8, 2017, 9:33 a.m.

Java

24 +0

95 +0

18 +0

GitHub
badger by dgraph-io

Fast key-value DB in Go.

created at Jan. 26, 2017, 5:09 a.m.

Go

239 +0

13,457 +12

1,151 +2

GitHub
awesome-memento by machawk1

A list of things related to software, literature, and other content for 🕣 Memento

created at Sept. 16, 2016, 1:33 a.m.

Unknown languages

8 +0

77 +0

8 +0

GitHub
HadoopConcatGz by helgeho

A Splitable Hadoop InputFormat for Concatenated GZIP Files and *.(w)arc.gz

created at Aug. 8, 2016, 1:36 p.m.

Java

2 +0

9 +0

3 +0

GitHub
heritrix-walkthrough by web-archive-group

None

created at June 1, 2016, 10:35 p.m.

Shell

6 +0

9 +0

1 +0

GitHub
wail by N0taN3rd

whale2 One-Click User Instigated Preservation

created at May 26, 2016, 4:52 a.m.

JavaScript

13 +0

120 +0

9 +0

GitHub
ipwb by oduwsdl

InterPlanetary Wayback: A distributed and persistent archive replay system using IPFS

created at March 4, 2016, 3:01 p.m.

Python

23 +0

591 +1

39 +0

GitHub