SPARK SQL - update MySql table using DataFrames and JDBC
2018-01-19 08:52
691 查看
I'm trying to insert and update some data on MySql using Spark SQL DataFrames and JDBC connection.I've succeeded to insert new data using the SaveMode.Append. Is there a way to update the data already existing in MySql Table from Spark SQL?My code to insert is:
It is not possible. As for now (Spark 1.6.0 / 2.2.0 SNAPSHOT) Spark
You can insert manually for example using
myDataFrame.write.mode(SaveMode.Append).jdbc(JDBCurl,mySqlTable,connectionProperties)If I change to SaveMode.Overwrite it deletes the full table and creates a new one, I'm looking for something like the "ON DUPLICATE KEY UPDATE" available in MySql
It is not possible. As for now (Spark 1.6.0 / 2.2.0 SNAPSHOT) Spark
DataFrameWritersupports only four writing modes:
SaveMode.Overwrite: overwrite the existing data.
SaveMode.Append: append the data.
SaveMode.Ignore: ignore the operation (i.e. no-op).
SaveMode.ErrorIfExists: default option, throw an exception at runtime.
You can insert manually for example using
mapPartitions(since you want an UPSERT operation should be idempotent and as such easy to implement), write to temporary table and execute upsert manually, or use triggers.In general achieving upsert behavior for batch operations and keeping decent performance is far from trivial. You have to remember that in general case there will be multiple concurrent transactions in place (one per each partition) so you have to ensure that there will no write conflicts (typically by using application specific partitioning) or provide appropriate recovery procedures. In practice it may be better to perform and batch writes to a temporary table and resolve upsert part directly in the database.
相关文章推荐
- 【MySQL笔记】解除输入的安全模式,Error Code: 1175. You are using safe update mode and you tried to update a table without a WHERE that uses a KEY column To disable safe mode, toggle the option in Preferences -> SQL Queries and reconnect.
- MySQL错误:Error Code: 1175. You are using safe update mode and you tried to update a table without a WHERE that uses a KEY column To disable safe mode, toggle the option in Preferences -> SQL easonjim
- mysql 更新sql脚本: you are using safe update mode and you tried to update a table
- Spark(六):SparkSQLAndDataFrames对结构化数据集与非结构化数据的处理
- 学习spark:五、Spark SQL, DataFrames and Datasets Guide
- Mysql: Table name is specified twice, both as a target for UPDATE and as a separate source for data
- Spark SQL and DataFrames
- Spark SQL, DataFrames and Datasets Guide
- Spark SQL,DataFrames and DataSets Guide官方文档翻译
- MySQL错误:You are using safe update mode and you tried to update a table without a WHERE that uses a K
- MySQL错误:You are using safe update mode and you tried to update a table without a WHERE that uses a K
- 【Spark】Spark SQL, DataFrames and Datasets Guide(翻译文,持续更新)
- Spark -9:Spark SQL, DataFrames and Datasets 编程指南
- MySQL更新表时 Error Code: 1175. You are using safe update mode and you tried to update a table......
- Apache Spark 2.2.0 中文文档 - Spark SQL, DataFrames and Datasets Guide | ApacheCN
- pyspark-Spark SQL, DataFrames and Datasets Guide
- Mysql Error Code: 1175. You are using safe update mode and you tried to update a table without a WHERE that uses a KEY column
- [SQL Server][FILESTREAM] -- Using INSERT, UPDATE and DELETE to manage SQL Server FILESTREAM Data
- Spark SQL and DataFrames Version 1.6
- Using Presto to combine data from Hive and MySQL in one SQL-like query