您的位置:首页 > 其它

Nutch1.7的deploy模式在伪分布式环境上报错

2014-02-13 17:35 106 查看
1.操作系统CentOS 6.5 x64

2.Hadoop平台为Cloudera CDH5 beta2,hadoop-2.2.0

3.开始操作:

    Nutch1.7已经编译成功,把seed.txt上传到HDFS的urls目录中,目标目录crawl不存在;

    在runtime/deploy下执行

hadoop  jar  apache-nutch-1.7.job  org.apache.nutch.crawl.Crawl  urls  -dir   crawl -depth 1  -topN 5
    正常情况会在crawl目录下生成抓取数据,但是此时会报错:
Exception in thread "main" java.lang.IllegalArgumentException: Wrong FS: hdfs://localhost/work/nutch/crawl/crawldb/918962832, expected: file:///
at org.apache.hadoop.fs.FileSystem.checkPath(FileSystem.java:644)
at org.apache.hadoop.fs.RawLocalFileSystem.pathToFile(RawLocalFileSystem.java:79)
at org.apache.hadoop.fs.RawLocalFileSystem.deprecatedGetFileStatus(RawLocalFileSystem.java:506)
at org.apache.hadoop.fs.RawLocalFileSystem.getFileLinkStatusInternal(RawLocalFileSystem.java:722)
at org.apache.hadoop.fs.RawLocalFileSystem.getFileStatus(RawLocalFileSystem.java:501)
at org.apache.hadoop.fs.FileSystem.isDirectory(FileSystem.java:1412)
at org.apache.hadoop.fs.ChecksumFileSystem.rename(ChecksumFileSystem.java:496)
at org.apache.nutch.crawl.CrawlDb.install(CrawlDb.java:159)
at org.apache.nutch.crawl.Injector.inject(Injector.java:297)
at org.apache.nutch.crawl.Crawl.run(Crawl.java:132)
at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:70)
at org.apache.nutch.crawl.Crawl.main(Crawl.java:55)
at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)
at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
at java.lang.reflect.Method.invoke(Method.java:606)
at org.apache.hadoop.util.RunJar.main(RunJar.java:212)

已知HDFS没有任何问题,自己编写的MR程序也可以运行,在第二个cloudera CDH4平台上也会报这个错
还未找到解决方案...
内容来自用户分享和网络整理,不保证内容的准确性,如有侵权内容,可联系管理员处理 点击这里给我发消息
标签: