Comments (4)
hi @sharpstill
Thank you for your request. We will consider this carefully because it is better to be a Chainer's feature, rather than ChainerMN's. In addition, we have limited resource and we don't have much experience on Hadoop/HDFS unfortunately.
Thanks!
from chainermn.
@keisukefukuda Thank you very much, HDFS can setup in only 2 machines to experiment.
By the way, tensorflow has already support HDFS and Yahoo release a tensorflow On spark version.
Maybe this will help you.
from chainermn.
I wonder in what @sharpstill meant by supporting HDFS file system. Even today it is possible to just access data in HDFS; by creating your own Dataset subclass that has list of files or chunks of HDFS file(s), and iterating over them at each get_example()
call. See our example on local file system and Chainer's documentation on dataset . With these tools it's already possible to make your job directly read contents from HDFS. Also see Wes McKinney's article on Python HDFS clients comparison for HDFS Python clients choice.
For further support such as data locality awareness and data access optimization, it needs various design decision and we'll be working on it.
from chainermn.
So we're testing some use cases on HDFS but still not feeling the need to special code for HDFS, as we can make own Dataset accessor like this, using libhdfs and PyArrow. For the time being I'm closing this issue, but feel free to reopen for further discussion.
from chainermn.
Related Issues (20)
- Don't inicialize global NCCL comm when HOT 2
- Checkpointer doesn't resume current learning rate HOT 8
- Adding allreduce for ndarray HOT 10
- mpirun doesn't exit when exception is thrown in some process HOT 7
- Asynchronous Allreduce HOT 2
- Handle list of dicts in MultiNodeIterator HOT 1
- would you please share hype parameters of GPUs=4 for resnet50 training with us ? HOT 23
- Expose `intra_size`, `inter_rank` and `inter_size` of communicators at readthedocs
- Provide functions for allreduce
- Manual selection for gpus in distributed training HOT 5
- CommunicatorBase.{scatter, allgather} is missing in the document
- Add `force_equal_length` flag to `scatter_dataset` method
- optimizer.setup() created by create_multi_node_optimizer returns an original optimizer HOT 2
- FP16 support HOT 1
- Forcing forkserver spawn earlier HOT 2
- When `in_size=None` is used in `Liner` and it is not used, an error occurs
- NCCL_ERROR_SYSTEM_ERROR: unhandled system error HOT 3
- CUDA streams usage HOT 6
- Non-Blocking Methodology on ChainerMN HOT 3
- Installation should do nothing but omit a warning.
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from chainermn.