Spark-1.0.0 SQL使用简介

最新推荐文章于 2023-03-17 11:49:41 发布

just-天之蓝

最新推荐文章于 2023-03-17 11:49:41 发布

阅读量490

点赞数

分类专栏： spark 文章标签： sql spark-sql

本文链接：https://blog.csdn.net/ZHAOLEI5911/article/details/74781881

版权

spark 专栏收录该内容

7 篇文章 0 订阅

订阅专栏

这里写图片描述

- 上传文件到HDFS
- 启动 sql

1.上传文件到HDFS

http://blog.csdn.net/zhaolei5911/article/details/64514726

2.启动 sql

spark1.0.0 中 sql 启动是直接在 spark-shell 启动后启动

val sqlContext = new org.apache.spark.sql.SQLContext(sc)
import sqlContext._

// Define the schema using a case class.
// Note: Case classes in Scala 2.10 can support only up to 22 fields. To work around this limit, 
// you can use custom classes that implement the Product interface.
case class Person(name: String, age: Int)

// Create an RDD of Person objects and register it as a table.
val people = sc.textFile("hdfs://hadoopmaster:8020/data/wordcount/people.txt").map(_.split(",")).map(p => Person(p(0), p(1).trim.toInt))
people.registerAsTable("people")

// SQL statements can be run by using the sql methods provided by sqlContext.
val teenagers = sql("SELECT name FROM people WHERE age >= 13 AND age <= 19")

// The results of SQL queries are SchemaRDDs and support all the normal RDD operations.
// The columns of a row in the result can be accessed by ordinal.
teenagers.map(t => "Name: " + t(0)).collect().foreach(println)