1.操作系统CentOS 6.5 x64
2.Hadoop平台为Cloudera CDH5 beta2,hadoop-2.2.0
3.开始操作:
Nutch1.7已经编译成功,把seed.txt上传到HDFS的urls目录中,目标目录crawl不存在;
在runtime/deploy下执行
hadoop jar apache-nutch-1.7.job org.apache.nutch.crawl.Crawl urls -dir crawl -depth 1 -topN 5
正常情况会在crawl目录下生成抓取数据,但是此时会报错:
Exception in thread "main" java.lang.IllegalArgumentException: Wrong FS: hdfs://localhost/work/nutch/crawl/crawldb/918962832, expected: file:///
at org.apache.hadoop.fs.FileSystem.checkPath(FileSystem.java:644)
at org.apache.hadoop.fs.RawLocalFileSystem.pathToFile(RawLocalFileSystem.java:79)
at org.apache.hadoop.fs.RawLocalFileSystem.deprecatedGetFileStatus(RawLocalFileSystem.java:506)
at org.apache.hadoop.fs.RawLocalFileSystem.getFileLinkStatusInternal(RawLocalFileSystem.java:722)
at org.apache.hadoop.fs.RawLocalFileSystem.getFileStatus(RawLocalFileSystem.java:501)
at org.apache.hadoop.fs.FileSystem.isDirectory(FileSystem.java:1412)
at org.apache.hadoop.fs.ChecksumFileSystem.rename(ChecksumFileSystem.java:496)
at org.apache.nutch.crawl.CrawlDb.install(CrawlDb.java:159)
at org.apache.nutch.crawl.Injector.inject(Injector.java:297)
at org.apache.nutch.crawl.Crawl.run(Crawl.java:132)
at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:70)
at org.apache.nutch.crawl.Crawl.main(Crawl.java:55)
at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)
at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
at java.lang.reflect.Method.invoke(Method.java:606)
at org.apache.hadoop.util.RunJar.main(RunJar.java:212)
已知HDFS没有任何问题,自己编写的MR程序也可以运行,在第二个cloudera CDH4平台上也会报这个错
还未找到解决方案...