bioinf-mcb / gisaid-scrapper Goto Github PK

View Code? Open in Web Editor NEW

41.0 5.0 16.0 44 KB

Scrapping tool for GISAID data regarding SARS-CoV-2

License: MIT License

Python 97.76% Dockerfile 1.69% Shell 0.55%

sars-cov-2 scraper selenium

gisaid-scrapper's Introduction

GISAID scrapper

Scrapping tool for GISAID data regarding SARS-CoV-2. You need an active account in order to use it.

Preparations

Install all requirements for the scrapper.

pip install -r requirements.txt

You need to download a for your operating system and place it in script's directory.

Your login and password can be provided in credentials.txt file in format:

login
password

Usage

usage: scrap.py [-h] [--username USERNAME] [--password PASSWORD]  
                          [--filename FILENAME] [--destination DESTINATION] 
                          [--headless [HEADLESS]] [--whole [WHOLE]]

optional arguments:
  -h, --help            show this help message and exit
  --username USERNAME, -u USERNAME
                        Username for GISAID
  --password PASSWORD, -p PASSWORD
                        Password for GISAID
  --filename FILENAME, -f FILENAME
                        Path to file with credentials (alternative, default:
                        credentials.txt)
  --destination DESTINATION, -d DESTINATION
                        Destination directory (default: fastas/)
  --headless [HEADLESS], -q [HEADLESS]
                        Headless mode (no browser window)
  --whole [WHOLE], -w [WHOLE]
                        Scrap whole genomes only

Example:

python3 scrap.py -u user -p pass -w

run the scrapper with username user and password pass, downloading only whole sequence data.

python3 scrap.py -w -q -d whole_genome

run the scrapper in headless mode with username and password read from credentials.txt, downloading only whole sequence data into whole_genome directory.

Result

The whole and partial genom sequences from GISAID will be downloaded into fastas/ directory. metadata.tsv file will also be created, containing following information for every sample:

Accession
Collection date
Location
Host
Additional location information
Gender
Patient age
Patient status
Specimen source
Additional host information
Outbreak
Last vaccinated
Treatment
Sequencing technology
Assembly method
Coverage
Comment
Length

as long as they were provided. You can interrupt the download and resume it later, the samples won't be downloaded twice. The tool has only been tested on windows 10.

Docker Image

It is also possible to run this scrapper in headless mode inside docker container. This allows to use it on any Operating System that is able to run Docker. Image created by Pawel Kulig and hosted on his DockerHub.

In this version, all parameters are provided via .env file -- login, password, destination, and whole genome flag.

Aside from gisaid_scrapper container Selenium contianer is used to operate in client server paradigm.

To run scrapper in container run:

docker-compose up

To run it detached add "-d" option.

To build Docker Image on your own run below command inside gisaid_scrapper directory:

docker build --tag name:tag .

geckodriver file inside gisaid_scrapper directory is required to perform this operation. See: https://github.com/mozilla/geckodriver/releases

gisaid-scrapper's People

Contributors

Stargazers

Watchers

Forkers

chwisteeng pedroelbanquero joannalange rpscruz primediscoveries pawelkulig1 llq0325 royhervel rintukutum askarlupka beansrowning areias qianqli faraz107 abuendia abulenciamiguel

gisaid-scrapper's Issues

Casing in fasta header ID

Currently, the FASTA header contains uppercase identifiers and the HCOV-19 prefix. The later one can be removed quite easy but patching the header ID into correct case is more complicated.
Can you correct it during write of the file to disk?

A workaround is to change the augur filter dictionary lookup (https://github.com/nextstrain/augur/blob/904ed6d6c154753cd2dd2d210768f2e4df1e8bb6/augur/filter.py#L135) to be case-insensitive.

CAPTCHA Issue

They made CAPTCHA in order to prevent users from Crawling information

FYI

Thanks

driver.execute_script doesn't work

I can run the code and it will open up GISAID and put in my login details. However, it cannot hit the login button or proceed any further.

Error message: selenium.common.exceptions.ElementClickInterceptedException: Message: Element <td class="yui-dt0-col-d yui-dt-col-d yui-dt-sortable yui-dt-resizeable"> is not clickable at point (x,y) because another element <div class="yui-dt-bd"> obscures it

Hi, I keep getting this error message when trying to download the fasta files with both your original script and this updated one. I see in the updated script that it includes a scrolling function, but I'm still running into issues with the program not being able to continue because it following file isn't visible. Can you assist? Thank you!

Program break and check the "high coverage only"

Thanks for your great program. It makes me relaxed.

Program sometimes breaks when click "next". I guess the page is not loaded yet. So I add "time.sleep(2)" in "def go_to_next_page" to solve this problem.

Suggest: add a option to check "high coverage only" and select host "human". That will be helpful for us. Thanks.

Sequence remove

GISAID usually remove some sequences, so some new upload sequence couldn't be download. Do you have any method to solve it? may confirm all downloaded sequences the time...

TimeoutException

I changed time timeout to 900 but it still doesn't work

python3 gisaid-scrapper/scrap.py --filename credentials.txt --destination gisaid
Traceback (most recent call last):
File "gisaid-scrapper/scrap.py", line 55, in
scrapper.login(login, passwd)
File "/space/s2/lenore/virus_evolution/gisaid-scrapper/gisaid_scrapper.py", line 83, in login
WebDriverWait(self.driver, 900).until(cond.staleness_of(login_box))
File "/home/lenore/bin/anaconda2/lib/python3.8/site-packages/selenium/webdriver/support/wait.py", line 80, in until
raise TimeoutException(message, screen, stacktrace)
selenium.common.exceptions.TimeoutException: Message:

Issue that the "Download button couldn't be found!"

Dear there,

I have a problem when running the script. I run the script: "python3 scrap.py -u ààà -p àààà -w", and got the error message about 30min later.

Traceback (most recent call last):
File "scrap.py", line 79, in
scrapper.download_packages('metadata_tsv')
File "/Users/qianqian/gisaid-scrapper/gisaid_scrapper.py", line 278, in download_packages
assert download is not None, "Download button couldn't be found!"
AssertionError: Download button couldn't be found!

Could you help to give some solution?

Thanks
Qianqian

time.sleep() too small

Couldn't download all genomes for two days (downloaded only 70%) because of errors. Every time script downloads some portion of genomes and raises an error. After increasing all time.sleep(x) to time.sleep(10) everywhere in the script i could finally download all genomes.

AMD Ryzen 3600x, 32 GB RAM, Ubuntu 18.04

Request to remove repository

Hello,

I'm relaying a request from GISAID to remove this repository. The use of these scrapers are negatively impacting the ability for GISAID to host this critical scientific data.

Thank you.

Add bulk download option

Would it be possible to add an option to download the bulk file, as when clicking "Download" on the bottom right of the browser, as well as the Acknowledgement Table?

Otherwise: Thanks for this truly useful package!

Element is stale

Hello,

thank you for making your scraper public.
Unfortunately it runs after downloading one sample in the issue below:

raise exception_class(message, screen, stacktrace)
selenium.common.exceptions.StaleElementReferenceException: Message: The element reference of is stale; either the element is no longer attached to the DOM, it is not in the current frame context, or the document has been refreshed

Arch Linux
Python 3.7.6 | packaged by conda-forge | (default, Mar 5 2020, 15:27:18)
geckodriver 0.26.0
Mozilla Firefox 74.0

FileNotFoundError: [Errno 2] No such file or directory

If '/' in the filename, the destination was confused.

Please check this issue.

Thank you in advance!

firefox driver not in path

Had to add path for firefoxdriver in ubuntu 18.04:
os.environ["PATH"] += os.pathsep + '/home/...'

Gisaid Scraper for Influenza

Were there any plans to expand this out to scraping information for influenza from GISAID? I haven't been able to find code for that, and I think this provides a template for that.