Save complete web page (incl css, images) using python/selenium

python selenium web-scraping web-crawler bioinformatics

As you noted, Selenium cannot interact with the browser's context menu to use Save as..., so instead to do so, you could use an external automation library like pyautogui.

pyautogui.hotkey('ctrl', 's')time.sleep(1)pyautogui.typewrite(SEQUENCE + '.html')pyautogui.hotkey('enter')

This code opens the Save as... window through its keyboard shortcut CTRL+S and then saves the webpage and its assets into the default downloads location by pressing enter. This code also names the file as the sequence in order to give it a unique name, though you could change this for your use case. If needed, you could additionally change the download location through some extra work with the tab and arrow keys.

Tested on Ubuntu 18.10; depending on your OS you may need to modify the key combination sent.

Full code, in which I also added conditional waits to improve speed:

import timefrom selenium import webdriverfrom selenium.webdriver.common.by import Byfrom selenium.webdriver.support.expected_conditions import visibility_of_element_locatedfrom selenium.webdriver.support.ui import WebDriverWaitimport pyautoguiURL = 'https://blast.ncbi.nlm.nih.gov/Blast.cgi?PROGRAM=blastx&PAGE_TYPE=BlastSearch&LINK_LOC=blasthome'SEQUENCE = 'CCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACA' #'GAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGAGAAGA'# open page with selenium# (first need to download Chrome webdriver, or a firefox webdriver, etc)driver = webdriver.Chrome()driver.get(URL)# enter sequence into the query field and hit 'blast' button to searchseq_query_field = driver.find_element_by_id("seq")seq_query_field.send_keys(SEQUENCE)blast_button = driver.find_element_by_id("b1")blast_button.click()# wait until results are loadedWebDriverWait(driver, 60).until(visibility_of_element_located((By.ID, 'grView')))# open 'Save as...' to save html and assetspyautogui.hotkey('ctrl', 's')time.sleep(1)pyautogui.typewrite(SEQUENCE + '.html')pyautogui.hotkey('enter')

python selenium web-scraping web-crawler bioinformatics

This is not a perfect solution, but it will get you most of what you need. You can replicate the behavior of "save as full web page (complete)" by parsing the html and downloading any loaded files (images, css, js, etc.) to their same relative path.

Most of the javascript won't work due to cross origin request blocking. But the content will look (mostly) the same.

This uses requests to save the loaded files, lxml to parse the html, and os for the path legwork.

from selenium import webdriverimport chromedriver_binaryfrom lxml import htmlimport requestsimport osdriver = webdriver.Chrome()URL = 'https://blast.ncbi.nlm.nih.gov/Blast.cgi?PROGRAM=blastx&PAGE_TYPE=BlastSearch&LINK_LOC=blasthome'SEQUENCE = 'CCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACAGCTCAAACACAAAGTTACCTAAACTATAGAAGGACA' base = 'https://blast.ncbi.nlm.nih.gov/'driver.get(URL)seq_query_field = driver.find_element_by_id("seq")seq_query_field.send_keys(SEQUENCE)blast_button = driver.find_element_by_id("b1")blast_button.click()content = driver.page_source# write the page contentos.mkdir('page')with open('page/page.html', 'w') as fp:    fp.write(content)# download the referenced files to the same path as in the htmlsess = requests.Session()sess.get(base)            # sets cookies# parse htmlh = html.fromstring(content)# get css/js files loaded in the headfor hr in h.xpath('head//@href'):    if not hr.startswith('http'):        local_path = 'page/' + hr        hr = base + hr    res = sess.get(hr)    if not os.path.exists(os.path.dirname(local_path)):        os.makedirs(os.path.dirname(local_path))    with open(local_path, 'wb') as fp:        fp.write(res.content)# get image/js files from the body.  skip anything loaded from outside sourcesfor src in h.xpath('//@src'):    if not src or src.startswith('http'):        continue    local_path = 'page/' + src    print(local_path)    src = base + src    res = sess.get(hr)    if not os.path.exists(os.path.dirname(local_path)):        os.makedirs(os.path.dirname(local_path))    with open(local_path, 'wb') as fp:        fp.write(res.content)

You should have a folder called page with a file called page.html in it with the content you are after.

python selenium web-scraping web-crawler bioinformatics

Inspired by FThompson's answer above, I came up with the following tool that can download full/complete html for a given page url (see: https://github.com/markfront/SinglePageFullHtml)

UPDATE - follow up with Max's suggestion, below are steps to use the tool:

Clone the project, then run maven to build:

$> git clone https://github.com/markfront/SinglePageFullHtml.git$> cd ~/git/SinglePageFullHtml$> mvn clean compile package

Find the generated jar file in target folder: SinglePageFullHtml-1.0-SNAPSHOT-jar-with-dependencies.jar
Run the jar in command line like:

$> java -jar .target/SinglePageFullHtml-1.0-SNAPSHOT-jar-with-dependencies.jar <page_url>

The result file name will have a prefix "FP, followed by the hashcode of the page url, with file extension ".html". It will be found in either folder "/tmp" (which you can get by System.getProperty("java.io.tmp"). If not, try find it in your home dir or System.getProperty("user.home") in Java).
The result file will be a big fat self-contained html file that includes everything (css, javascript, images, etc.) referred to by the original html source.

CodeHunter

Save complete web page (incl css, images) using python/selenium

Recent Posts

How can I color dots in a xy scatterplot according to column value?

How to update a claim in ASP.NET Identity?

What does {0} mean when initializing an object?

Accessing members of items in a JSONArray with Java

How to log SQL statements in Spring Boot?

Powershell Get-WebSite name parameter is ignored

How to detect scroll to bottom of html element

Java synchronized method

How to test controllers with CodeIgniter?

Detect Visual Composer

Matplotlib: Specify format of floats for tick labels

Rails join a list of strings with commas and "and" before the last