Posts

Showing posts with the label web-scraping

VBA Webscrape not picking up elmenents; pick up frames/tables?

Image
VBA Webscrape not picking up elmenents; pick up frames/tables? Tried asking this question. Didn't get many answers. Can't install things onto my work computer. https://stackoverflow.com/questions/29805065/vba-webscrape-not-picking-up-elements Want to scrape a morningstar page into Excel with the code below. Problem is, it doesn't feed any real elements/data back. I actually just want the Dividend and cap gain distribution table really from that link I put into my_Page. This is usually easiest way, but an entire page scrape way, AND Excel-->Data-->From Web DON'T work. I've tried to use get elements by tag name and class before, but I failed at being able to do it in this case.This might be the way to go... Once again, just want that Dividend and Cap Gain distribution table. Not seeing any results in via the Debug.print Working code below, just need to parse into excel. Updated attempt below: Sub Macro1() Dim IE As New InternetExplorer IE.Visible = True ...

BS4/Python3 can't open other href while scrapping on google

BS4/Python3 can't open other href while scrapping on google My job is with a startup and they're calling some businesses but they're buying the contacts. So I had the idea to scrape them from Google, like some hotels, etc... I can already get the link that opens the Googlemaps with lots of companies but can't take the information inside this link because the program crashes. import json from bs4 import BeautifulSoup as bs from collections import namedtuple from pprint import pprint from requests import get import requests def remove_escape(s): return ' '.join(s.split()) def get_jobs(url): vagas = get(url, headers=headers) vagas_page = bs(vagas.text, 'html.parser') boxes = vagas_page.find_all('div', {'class': 'idQ6DBVUh1_8- ptqfrjbX76M'}) for box in boxes: titulo = box.find('span', {'class': 'ellip'}).text ...

NoneType error during multi-page scrape

NoneType error during multi-page scrape I'm working on a web scraper and am close to getting what I need, but I can't figure out why I'm getting a NoneType error all of a sudden after it finishes scraping the fourth page (of 204). Here's my code: script_path = os.path.dirname(os.path.realpath(__file__)) driver = webdriver.PhantomJS(executable_path="/usr/local/bin/bin/phantomjs", service_args=['--ignore-ssl-errors=true', '--ssl-protocol=any']) case_list = #this function launches the headless browser and gets us to the first page of results, which we'll scrape using main def search(): driver.get('https://www.courts.mo.gov/casenet/cases/nameSearch.do') if 'Service Unavailable' in driver.page_source: log('Casenet website seems to be down. Receiving "service unavailable"') driver.quit() gc.collect() return False time.sleep(2) court = Select(driver.find_element_by_i...

Excel VBA Macro: Scraping data from site table that spans multiple pages

Excel VBA Macro: Scraping data from site table that spans multiple pages Thanks in advance for the help. I'm running Windows 8.1, I have the latest IE / Chrome browsers, and the latest Excel. I'm trying to write an Excel Macro that pulls data from StackOverflow (https://stackoverflow.com/tags). Specifically, I'm trying to pull the date (that the macro is run), the tag names, the # of tags, and the brief description of what the tag is. I have it working for the first page of the table, but not for the rest (there are 1132 pages at the moment). Right now, it overwrites the data everytime I run the macro, and I'm not sure how to make it look for the next empty cell before running.. Lastly, I'm trying to make it run automatically once per week. I'd much appreciate any help here. Problems are: Code (so far) is below. Thanks! Enum READYSTATE READYSTATE_UNINITIALIZED = 0 READYSTATE_LOADING = 1 READYSTATE_LOADED = 2 READYSTATE_INTERACTIVE = 3 READYSTATE_COMPLETE = 4 ...

beautifulsoup multiple keyword from csv file

beautifulsoup multiple keyword from csv file I have a csv file with 2 column A and B, and I want scrap all the file with beautifulsoup The url is composed like this : http://.../search?info=A&who=B how to create a loop? my code from bs4 import BeautifulSoup import requests import json import csv with open('input.csv') as csvfile: reader = csv.reader(csvfile) for row in reader: url = ".../search?info={}&who={}".format(row[0], row[1]) response = requests.get(url) html = response.content soup = BeautifulSoup(html, "html5lib") for p in soup.find_all(class_="crd"): b = p.find(class_="info") if b['data-info'] is not None: j = json.loads(b['data-info']) data= p.h2.a.string Why do you need a loop? – Mad Physicist Jun 29 at 17:41 ...

Attribute Error: 'None Type' object has no attribute 'get_text'

Attribute Error: 'None Type' object has no attribute 'get_text' I have tried this program as I am scrapping data from Amazon but this program giving me the error. instead of get_text I also tried extract() and only strip () also they all are giving Attribute error. Now please help me what should I do? import urllib.request from bs4 import BeautifulSoup import pymysql.cursors a = input ('enter the item to be searched :') a = a.replace(" ","") html = urllib.request.urlopen("https://www.amazon.in/s/ref=nb_sb_noss?url=search-alias%3Daps&field-keywords="+a) bsObj = BeautifulSoup(html,'lxml') recordList = bsObj.findAll('a', class_='a-link-normal a-text-normal') connection = pymysql.connect(host='localhost', user='root', password='', db='shopping', charset='utf8mb4', ...

Output for scrapped data with Text value in CSV file

Output for scrapped data with Text value in CSV file I am new to web scrapping using Python and need help on extracting Sub category name (Title) and the Page Title (Main Category Header) with the URLs that is being scrapped with my Python code. I tried .text with beautifulsoup but I think there could be better option to do this task as I am getting error and no output once used. Help would be appreciated. Please take a look at the code and help with Output stored in csv file with URL t Sub category title t Main Category header. Example: Subcategory URL Required: http://www.medicalexpo.com/medical-manufacturer/neonatal-incubator-2963.html Neonatal incubators Pediatrics http://www.medicalexpo.com/medical-manufacturer/infant-radiant-warmer-13522.html Infant radiant warmers Pediatrics http://www.medicalexpo.com/medical-manufacturer/infant-phototherapy-lamp-44327.html Infant phototherapy lamps Pediatrics Something like this Code: from bs4 import Bea...