您现在的位置是：首页 > 后端

当前栏目

[python][spark]wholeTextFiles 读入多个文件的例子

Python 文件 Spark 多个例子读入

2023-09-11 14:20:28 时间

$pwd

/home/training/mydir

$cat file1.json

{
"firstName":"Fred",
"lastName":"Flintstone",
"userid":"123"
}

$cat file2.json

{
"firstName":"Barney",
"lastName":"Rubble",
"userid":"123"
}

[training@localhost ~]$ hdfs dfs -put /home/training/mydir
[training@localhost ~]$
[training@localhost ~]$ hdfs dfs -ls
Found 4 items
drwxrwxrwx - training supergroup 0 2017-09-23 19:26 .sparkStaging
-rw-rw-rw- 1 training supergroup 48 2017-09-25 05:31 cats.txt
drwxrwxrwx - training supergroup 0 2017-09-25 15:39 mydir ***
-rw-rw-rw- 1 training supergroup 34 2017-09-23 06:16 test.txt
[training@localhost ~]$

myrdd1 = sc.wholeTextFiles("mydir")

myrdd1.count()
Out[32]: 2

In [35]: myrdd1.take(2)

Out[35]:
[(u'hdfs://localhost:8020/user/training/mydir/file1.json',
u'{\n "firstName":"Fred",\n "lastName":"Flintstone",\n "userid":"123"\n}\n'),
(u'hdfs://localhost:8020/user/training/mydir/file2.json',
u'{\n "firstName":"Barney",\n "lastName":"Rubble",\n "userid":"456"\n}\n')]

猜你喜欢

安洵杯2022 Web Writeup
日期选择器
如何在JSP里自定义标签
CentOS6.5 切换图形界面与命令行界面
spss pro网络挑战赛A题：人群疏散模拟代码
EMS SQL Manager for PostgreSQL v6.4 Crack
RecyclerView 和 ListView 使用对比分析
why I cannot get any search result from P8F
SAP UI5 ObjectPageLayout 控件使用方法分享
Mybatis+mysql动态分页查询数据案例——条件类（HouseCondition）
一文秒懂串口、COM口、TTL、RS-232、RS-485区别
第十六周oj刷题——Problem J: 填空题：静态成员---计算学生个数

相关主题

python读xml文件
Python Tornado
Python 文件操作
Python-Python入门
python 文件处理
Python爬虫笔记
SQLSERVER文件和文件组
python：文件
python redis 操作

zl程序教程

当前栏目

[python][spark]wholeTextFiles 读入多个文件的例子

相关文章