在Postgres中进行批量插入的最快方法是什么?

我需要通过编程的方式将数千万条记录插入Postgres数据库。目前，我在一个查询中执行了数千条插入语句。

有没有更好的方法来做到这一点，一些我不知道的批量插入语句?

当前回答

我刚刚遇到了这个问题，建议将csvsql(发行版)批量导入到Postgres。要执行批量插入，只需创建b，然后使用csvsql，它连接到数据库，并为整个csv文件夹创建单独的表。

$ createdb test 
$ csvsql --db postgresql:///test --insert examples/*.csv

2015-08-13 15:08:49

其他回答

你可以使用COPY表TO…使用二进制，它“比文本和CSV格式略快”。只有当您有数百万行要插入，并且您对二进制数据感到满意时才这样做。

下面是一个使用psycopg2和二进制输入的Python食谱示例。

2011-11-17 09:33:08

带有数组的UNNEST函数可以与多行VALUES语法一起使用。我认为这个方法比使用COPY慢，但它对我在psycopg和python (python列表传递给游标)的工作中很有用。execute变成pg ARRAY

INSERT INTO tablename (fieldname1, fieldname2, fieldname3)
VALUES (
    UNNEST(ARRAY[1, 2, 3]), 
    UNNEST(ARRAY[100, 200, 300]), 
    UNNEST(ARRAY['a', 'b', 'c'])
);

没有值使用subselect额外的存在性检查:

INSERT INTO tablename (fieldname1, fieldname2, fieldname3)
SELECT * FROM (
    SELECT UNNEST(ARRAY[1, 2, 3]), 
           UNNEST(ARRAY[100, 200, 300]), 
           UNNEST(ARRAY['a', 'b', 'c'])
) AS temptable
WHERE NOT EXISTS (
    SELECT 1 FROM tablename tt
    WHERE tt.fieldname1=temptable.fieldname1
);

同样的语法用于批量更新:

UPDATE tablename
SET fieldname1=temptable.data
FROM (
    SELECT UNNEST(ARRAY[1,2]) AS id,
           UNNEST(ARRAY['a', 'b']) AS data
) AS temptable
WHERE tablename.id=temptable.id;

2015-06-26 10:37:57

加快速度的一种方法是在一个事务中显式地执行多个插入或复制(比如1000个)。Postgres的默认行为是在每条语句之后提交，因此通过批处理提交，可以避免一些开销。正如Daniel回答中的指南所说，您可能必须禁用自动提交才能工作。还要注意底部的注释，该注释建议将wal_buffers的大小增加到16mb也可能有所帮助。

2009-04-17 04:06:48

May be I'm late already. But, there is a Java library called pgbulkinsert by Bytefish. Me and my team were able to bulk insert 1 Million records in 15 seconds. Of course, there were some other operations that we performed like, reading 1M+ records from a file sitting on Minio, do couple of processing on the top of 1M+ records, filter down records if duplicates, and then finally insert 1M records into the Postgres Database. And all these processes were completed within 15 seconds. I don't remember exactly how much time it took to do the DB operation, but I think it was around less then 5 seconds. Find more details from https://www.bytefish.de/blog/pgbulkinsert_bulkprocessor.html

2021-07-29 17:23:02

$ createdb test 
$ csvsql --db postgresql:///test --insert examples/*.csv

2015-08-13 15:08:49

在Postgres中进行批量插入的最快方法是什么?

推荐文章

最新文章

标签