warcbase

Warcbase is an open-source platform for managing analyzing web archives

  • 所有者: lintool/warcbase
  • 平台:
  • 许可证:
  • 分类:
  • 主题:
  • 喜欢:
    0
      比较:

Github星跟踪图

Warcbase

Warcbase is an open-source platform for managing web archives built on Hadoop and HBase. The platform provides a flexible data model for storing and managing raw content as well as metadata and extracted knowledge. Tight integration with Hadoop provides powerful tools for analytics and data processing via Spark.

Bad news: Warcbase is defunct and no longer under active development!

Good news: In June 2017, the University of Waterloo and York University were awarded a grant from the Andrew W. Mellon Foundation to build the next generation of tools that will make historical internet content accessible to scholars. Warcbase serves as the foundation for the ArchivesUnleashed Toolkit!

If you're interested in reading about the development of Warcbase, check out this article:

Jimmy Lin, Ian Milligan, Jeremy Wiebe, and Alice Zhou. Warcbase: Scalable Analytics Infrastructure for Exploring Web Archives. ACM Journal on Computing and Cultural Heritage, 10(4), Article 22, 2017.

License

Licensed under the Apache License, Version 2.0.

Acknowledgments

This work has been supported in part by the U.S. National Science Foundation, the Natural Sciences and Engineering Research Council of Canada, the Social Sciences and Humanities Research Council of Canada, the Ontario Ministry of Research and Innovation's Early Researcher Award program, and the Mellon Foundation (via Columbia University). Any opinions, findings, and conclusions or recommendations expressed are those of the researchers and do not necessarily reflect the views of the sponsors.

主要指标

概览
名称与所有者lintool/warcbase
主编程语言Java
编程语言Java (语言数: 6)
平台
许可证
所有者活动
创建于2013-07-13 19:13:19
推送于2017-12-08 02:51:58
最后一次提交2017-09-13 22:26:34
发布数0
用户参与
星数162
关注者数23
派生数47
提交数714
已启用问题?
问题数229
打开的问题数34
拉请求数23
打开的拉请求数4
关闭的拉请求数6
项目设置
已启用Wiki?
已存档?
是复刻?
已锁定?
是镜像?
是私有?