Comments (7)
from datacomp.
I am sorry I am working in campus network, for security-related reasons I can't change the DNS, is there alternative solution for DNS?
from datacomp.
from datacomp.
from datacomp.
If I use a proxy to download the data, should I set some parameters in img2dataset? And I think it's strange that the downloaded medium_data/shards takes up 660G, and the complete size is 750G(in README), why the shell show
"total - success: 0.134 - failed to download: 0.865 - failed to resize: 0.001"
shards takes 660/750=0.88, I'm glad to see your insights.
from datacomp.
Hi Team,
We are getting success rate of +- 86% (downloading the small, 12 million images, dataset):
15it [44:26, 19.65s/it]worker - success: 0.867 - failed to download: 0.128 - failed to resize: 0.005 - images per sec: 4 - count: 10000 total - success: 0.865 - failed to download: 0.130 - failed to resize: 0.005 - images per sec: 56 - count: 150000
16it [44:49, 20.77s/it]worker - success: 0.860 - failed to download: 0.134 - failed to resize: 0.006 - images per sec: 4 - count: 10000 total - success: 0.864 - failed to download: 0.131 - failed to resize: 0.005 - images per sec: 60 - count: 160000
We are using knot dns resolver (8 instances), using the default 16 processes, and using just 16 threads (instead of the default 128). We did this to slow down the dns resolve requests. We use an e2-highcpu-32 (Efficient Instance, 32 vCPUs, 32 GB RAM) instance on GCP.
86% is the best we could get, but due to the low number of threads, the downloading is very slow. Can you share any advice or tips to improve speed and/or success rate? Many thanks!
from datacomp.
from datacomp.
Related Issues (20)
- 14% of SHA256 hashes not matching HOT 32
- the normal success rate and downloading speed? HOT 1
- `zeroshot_templates` split error for FairFace / UTKFace HOT 9
- Deduplication against evaluation sets HOT 1
- Remove CSAM, if present HOT 2
- Metadata for datacomp-large text-based filter HOT 1
- Pretraining dataset HOT 1
- Training log HOT 1
- Frequency of Leaderboard Updates HOT 1
- About update metadata with the corresponding image sample in shards HOT 2
- ModuleNotFoundError: No module named 'training' HOT 2
- Availability of npy indices for large pool
- Average caption length for CommonPool HOT 1
- Downloading Commonpool XLarge
- ImageNet 21k based filtered dataset HOT 1
- Invalid files for Datacomp1B
- Problems in run train.py HOT 3
- Metadata downloading fails and no way to resume the download
- Redundant labels in iWILDCAM eval data
- Label Errors in ImageNet-O Eval Set
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from datacomp.