-
Notifications
You must be signed in to change notification settings - Fork 2
next steps (10/04/2024) #19
Copy link
Copy link
Open
Description
amchagas
opened on Apr 10, 2024
Issue body actions
- Push latest code (Andre)
- learn a solution for data storage (datalad, git annex, or git lfs) (Andre and Ale)
- We chose Datalad as a solution. See:
- https://handbook.datalad.org/en/latest/_images/datalad-cheatsheet_p1.png
- https://handbook.datalad.org/en/latest/basics/101-105-install.html
- We chose Datalad as a solution. See:
- sort out pdf download storage for zotero (Andre)
- figure out issues with "requests" and see if "curl" implementation works better for getting paper metadata (Ale).
- Try these potential solutions:
- https://requests.readthedocs.io/en/latest/user/advanced/#timeouts
- https://docs.python.org/3/library/urllib.request.html#urllib.request.urlopen
- Try these potential solutions:
- figure out based on corpus size the number of papers that need to be analysed (Ale)
- re run current analysis (Ale)
- Check size of HardwareX (517) and Journal of Open Hardware size (41)
- Once zotero metadata and storage are sorted, re-run the pdf miner code
Reactions are currently unavailable