patroni

mirror of https://github.com/outbackdingo/patroni.git synced 2026-01-28 02:20:04 +00:00

Author	SHA1	Message	Date
Dmitry Dolgov	11f7ceb521	Do not check types of standby_cluster configuration (#924 ) Simply allow valid keys	2019-01-14 14:16:15 +01:00
Alexander Kukushkin	f1d7ccf36e	Make sure we refresh session at least once per HA loop (#880 ) Fixes https://github.com/zalando/patroni/issues/879	2018-12-03 16:35:14 +01:00
Alexander Kukushkin	9bf074acfb	Compatibility with python3 (#883 ) Change of `loop_wait` was causing Patroni to disconnect from zookeeper and never reconnect back. The error was happening only with python3 due to a difference in implementation of `select.select` function.	2018-11-30 11:40:34 +01:00
Alexander Kukushkin	fb01aaebc5	Compatibility with kazoo-2.6.0 (#872 ) Recently 2.6.0 was release which changes the way how create_connection method is called. Before it was passing two arguments, and in the new version all argument names are specified explicitly.	2018-11-19 14:26:20 +01:00
Alexander Kukushkin	0f666e69f3	Prefix system tables, views and functions with pg_catalog (#845 ) and implement missing unit tests	2018-11-01 16:17:40 +01:00
Alexander Kukushkin	2efd97baab	Permanent replication slots (#819 ) Permanent replication slots are preserved on failover/switchover, that is Patroni on the new primary will create configured replication slots right after doing promote. Slots could be configured with the help of `patronictl edit-config`. The initial configuration could be also done in the `bootstrap.dcs` ```yaml slots: permanent_physical_1: type: physical permanent_logical_1: type: logical database: foo plugin: pgoutput ``` It is the responsibility of the operator to make sure that there are no clashes in names between replication slots automatically created by Patroni for members and permanent replication slots. Closes https://github.com/zalando/patroni/issues/656	2018-10-31 11:37:42 +01:00
Alexander Kukushkin	f70edefc65	A few bugfixes in the "standby cluster" workflow (#823 ) * Always run `pg_rewind` against the remote master * Always use the remote master as the source when "recovering" stopped standby leader * Use remote master as the source when "recovering" the node in the unhealthy cluster * Use the local dbname as the fallback when doing `pg_rewind` from the remote master * `no_replication_slot` is the allowed key in the `RemoteMember` object * Make it possible to "bootstrap" the new `standby_cluster` with existing (and valid) data directory. There is one prerequisite though, there should be no `patroni.dynamic.json` file in it!	2018-10-09 13:30:48 +02:00
Alexander Kukushkin	76d1b4cfd8	Minor fixes (#808 ) * Use `shutil.move` instead of `os.replace`, which is available only from 3.3 * Introduce standby-leader health-check and consul service * Improve unit tests, some lines were not covered * rename `assertEquals` -> `assertEqual`, due to deprecation warning	2018-09-19 16:32:33 +02:00
Pavel Kirillov	2e9cb412e4	Register service in consul (#802 ) Кegister service 'scope_name' with tag 'master' or 'replica' example with scope 'pgsql-pgpi' ```[root@pgpi1 ~]# host -t SRV pgsql-pgpi.service.consul. 127.0.0.1 Using domain server: Name: 127.0.0.1 Address: 127.0.0.1#53 Aliases: pgsql-pgpi.service.consul has SRV record 1 1 5432 pgpi1.node.dc.consul. pgsql-pgpi.service.consul has SRV record 1 1 5432 pgpi2.node.dc.consul. [root@pgpi1 ~]# host -t SRV master.pgsql-pgpi.service.consul. 127.0.0.1 Using domain server: Name: 127.0.0.1 Address: 127.0.0.1#53 Aliases: master.pgsql-pgpi.service.consul has SRV record 1 1 5432 pgpi2.node.dc.consul. [root@pgpi1 ~]# host -t SRV replica.pgsql-pgpi.service.consul. 127.0.0.1 Using domain server: Name: 127.0.0.1 Address: 127.0.0.1#53 Aliases: replica.pgsql-pgpi.service.consul has SRV record 1 1 5432 pgpi1.node.dc.consul.``` Fixes: https://github.com/zalando/patroni/issues/771	2018-09-07 15:17:56 +02:00
Dmitry Dolgov	dd7c3c349f	[WIP] Standby cluster implementation (#679 ) Implementation of "standby cluster" described in #657. Standby cluster consists of a "standby leader", that replicates from a "remote master" (which is not a part of current patroni cluster and can be anywhere), and cascade replicas, that replicate from the corresponding standby leader. "Standby leader" behaves pretty much like a regular leader, which means that it holds a leader lock in DSC, in case if disappears there will be an election of a new "standby leader". One can define such a cluster using the section "standby_cluster" in patroni config file. This section provides parameters for standby cluster, that will be applied only once during bootstrap and can be changed only through DSC.	2018-09-07 10:10:56 +02:00
Alexander Kukushkin	4ca8a6e506	Make retries of calls to DCS consistent across implementations (#805 ) in addition to that do a small refactoring of zookeeper and consul and try to improve the stability of AT	2018-09-06 08:37:26 +02:00
wilfriedroset	0136f252ab	Add patronictl -k/--insecure flag and suport for restapi cert (#790 ) Fixes https://github.com/zalando/patroni/issues/785	2018-08-29 16:08:13 +02:00
Alexander Kukushkin	90cf930036	Refactor REST API health-checks (#779 ) Make it more readable and easy to understand. Mostly it is needed to implement https://github.com/zalando/patroni/issues/772	2018-08-29 11:35:22 +02:00
Alexander Kukushkin	87e9aab04c	Improve tests (#778 ) * Implement missing unit-tests * Add acceptance tests for ISSUE #776 * Update list of classifiers, keywords and authors	2018-08-29 11:29:37 +02:00
Alexander Kukushkin	0c1ae6fbeb	Respond 200 to master health-check only if update_lock was successful (#713 ) If Patroni gets partitioned it starts receiving stale information from DCS. We can't use this information to determine that we have the leader key. Instead, we will record in Ha object the actual state of acquire/update lock and report as a leader only if it was successful. P.S. despite responding with 200 on `GET /master` postgres was still running read-only.	2018-08-03 17:00:01 +02:00
Alexander Kukushkin	8a3b78ca7b	Rest api thread can raise an exception during shutdown (#711 ) catch it and report	2018-06-14 13:17:50 +02:00
Dmitry Dolgov	f0d23b0b14	Merge pull request #706 from zalando/feature/rename-create-replica-method Rename create_replica_method to create_replica_methods	2018-06-12 14:16:54 +02:00
Alexander Kukushkin	aadd39b0a4	Do crash recovery only when we sure that postgres was running as master (#707 ) pg_controldata reports in this case: * 'in production' * 'shutting down' * 'in crash recovery'	2018-06-12 14:09:09 +02:00
Henning Jacobs	2537147810	#694 handle configuration error (#695 ) It is possible to change a lot of parameters in runtime (including `restapi.listen`) by updating Patroni config file and sending SIGHUP to Patroni process. If something was misconfigured it was throwing a weird exception and breaking `restapi` thread. This PR improves friendliness of error message and avoids breaking of `restapi`.	2018-06-12 14:08:38 +02:00
Alexander Kukushkin	e939304001	Take and apply some parameters from controldata when starting as replica (#703 ) * Take and apply some parameters from controldata when starting as replica https://www.postgresql.org/docs/10/static/hot-standby.html#HOT-STANDBY-ADMIN There is set of parameters which value on the replica must be not smaller than on the primary, otherwise replica will refuse to start: * max_connections * max_prepared_transactions * max_locks_per_transaction * max_worker_processes It might happen that values of these parameters in the global configuration are not set high enough, what makes impossible to start a replica without human intervention. Usually it happens when we bootstrap a new cluster from the basebackup. As a solution to this problem we will take values of above parameters from the pg_controldata output and in case if the values in the global configuration are not high enough, apply values taken from pg_controldata and set `pending_restart` flag.	2018-06-12 14:04:32 +02:00
Alexander Kukushkin	e405e4e03c	Workaround to sporadic unit-test failures (#696 ) Fixes https://github.com/zalando/patroni/issues/691	2018-06-12 14:00:10 +02:00
erthalion	d037aa8afd	Rename create_replica_method to create_replica_methods To make it clear that it's actually an array	2018-06-12 11:33:13 +02:00
Alexander Kukushkin	856552bd61	Sync replication slots and verify sysid after coming out of pause (#678 ) Fixes https://github.com/zalando/patroni/issues/568 and https://github.com/zalando/patroni/issues/674	2018-05-18 12:18:49 +02:00
Oleksii Kliukin	4ce539ba1b	Allow options to the basebackup built-in method. (#604 ) Options should be specified in the basebackup section, which is optional.	2018-05-18 12:18:35 +02:00
Oleksii Kliukin	1043376e6b	Do not exit when encountering invalid system ID. (#669 ) Do not exit when the cluster system ID is empty or the one that doesn't pass the validation check. In that case, the cluster most likely needs a reinit; mention it in the result message. Avoid terminating Patroni, as otherwise reinit cannot happen.	2018-05-18 11:48:15 +02:00
Alexander Kukushkin	ed479fe585	Don't demote master if failed to update leader key in pause (#668 ) Fixes https://github.com/zalando/patroni/issues/659	2018-05-18 11:19:56 +02:00
Alexander Kukushkin	5ce18a8045	Improve protection of DCS being accidentally wiped (#680 ) We already have a lot of logic in place to prevent failover in such case and restore all keys, but an accidental removal of `/config` key was effectively switching off pause mode for 1 cycle of HA loop.	2018-05-18 11:18:58 +02:00
Alexander Kukushkin	5296336f4a	BUGFIX: postmaster start can fail if pid from postmaster.pid is alive (#681 ) Upon start postmaster process performs various safety checks if there is a postmaster.pid file in the data directory. Although Patroni already detected that the running process corresponding to the postmaster.pid is not a postmaster, the new postmaster might fail to start, because it thinks that postmaster.pid is already locked. Important!!! Unlink of postmaster.pid isn't an option in this case, because it has a lot of nasty race conditions. Luckily there is a workaround to this problem, we can pass the pid from postmaster.pid in the `PG_GRANDPARENT_PID` environment variable and postmaster will ignore it. More likely to hit such problem if you run Patroni and postgres in the docker container.	2018-05-18 11:18:27 +02:00
Alexander Kukushkin	84f29caf92	Fix race condition in poll_failover_result (#658 ) It didn't affect directly neither failover nor switchover, but in some rare cases it was reporting it as a success too early, when the former leader released the lock: `Failed over to "None" instead of "desired-node"` In addition to that this commit improves logs and status messages by differentiating between failover and switchover.	2018-04-16 17:45:05 +02:00
Alexander Kukushkin	d78790b194	Abort start if attaching to running postgres and cluster not initiazlied (#661 ) Patroni can attach itself to an already running PostgreSQL instance. If that is the first instance "seen" in the given cluster, Patroni for that instance will create the initialize key, grab the leader key and, if the instance is running a replica, promote. Because of this behavior, when a cluster with a master and one or more replicas gets Patroni for each node, it is imperative to start running Patroni on the master node before getting to the replicas. This commit changes such weird behavior and will abort Patroni start if there is no initialize key in DCS and postgres is running as a replica. Closes https://github.com/zalando/patroni/issues/655	2018-04-16 17:32:26 +02:00
Alexander Kukushkin	3afd26101b	Single user mode was waiting for user input and never finish (#634 ) Regression was introduced in https://github.com/zalando/patroni/pull/576	2018-03-02 22:22:43 +01:00
Alexander Kukushkin	c04e7a1798	Write bootstrap.pg_hba into a pg_hba.conf after custom bootstrap (#632 ) Fixes https://github.com/zalando/patroni/issues/631	2018-02-26 18:48:56 +01:00
Alexander Kukushkin	89a11fed07	Don't rediscover etcd cluster topology when watch timed out (#630 ) but switch to the next node if it is possible. Fixes https://github.com/zalando/patroni/issues/628	2018-02-26 18:48:30 +01:00
Alexander Kukushkin	dd1500b4dc	Handle exceptions raised from psutil (#610 ) Process.cmdline can raise `NoSuchProcess` or `AccessDenied` Process.children can raise `AccessDenied` Fixes https://github.com/zalando/patroni/issues/609	2018-01-30 16:26:39 +01:00
Alexander Kukushkin	a0c8491abb	Don't swallow silently all errors from k8s API (#611 ) Output exception trace to the logs when http status code == 403, something is wrong with permissions. When http status code == 409 -- such error could be ignored, because object probably was created or updated by another process. For all other http status codes it will also produce stack traces. I hope it will help to debug issues similar to the https://github.com/zalando/patroni/issues/606	2018-01-26 09:57:17 +01:00
Alexander Kukushkin	a1e5c8e1cb	A few iprovements in patronictl (#601 ) * make switchover work with an old patroni * exclude leader from candidates when interactively running failover	2018-01-17 15:33:08 +01:00
Alexander Kukushkin	5668367181	Implement '/sync' and `/async` endpoints (#578 ) They will respond with http status code 200 only when the node is running as a synchronous or asynchronous replica. Fixes https://github.com/zalando/patroni/issues/189 Fixes https://github.com/zalando/patroni/issues/415	2018-01-05 15:28:40 +01:00
Alexander Kukushkin	03c2a85d23	Expose current timeline in DCS and via API (#591 ) It is very easy to get current timeline on the master by executing ```sql SELECT ('x' \|\| SUBSTR(pg_walfile_name(pg_current_wal_lsn()), 1, 8))::bit(32)::int ``` Unfortunately the same method doesn't work when postgres is_in_recovery. Therefore we will use replication connection for that on the replicas. In order to avoid opening and closing replication connection on every HA loop we will cache the result if its value matches with the timeline of the master. Also this PR introduces a new key in DCS: `/history`. It will contain a json serialized object with timeline history in a format similar to the usual history files. The differences are: * Second column is the absolute wal position in bytes, instead of LSN * Optionally there might be a fourth column - timestamp, (mtime of history file)	2018-01-05 15:25:56 +01:00
Alexander Kukushkin	18786464a1	Rename failover to switchover and make new failover work without leader (#588 ) In addition to that implement /switchover endpoint as an alias to /failover endpoint and implement more checks like: * candidate must be provided for a failover * switchover can't be scheduled in a pause state * and so on Fixes https://github.com/zalando/patroni/issues/585 Fixes https://github.com/zalando/patroni/issues/520	2018-01-05 15:17:56 +01:00
Alexander Kukushkin	3a96ffa718	Expose pause state of every member to DCS and via REST (#592 ) and implement patronictl pause\|resume --wait on top of that Fixes https://github.com/zalando/patroni/issues/349	2018-01-05 15:16:45 +01:00
Alexander Kukushkin	6b01d2787f	More improvements in patronictl (#590 ) Make specifying cluster_name optional for some more commands. If it is not specified, it's value would be taken from config file.	2018-01-04 12:26:13 +01:00
Alexander Kukushkin	2b8618b027	Minimize amount of SELECTS issued by Patroni on every loop (#584 ) Every iteration of HA loop Patroni needs to call pg_is_in_recovery() and calcualte absolute wal_position. It was doing two separate SELECT statements for that. In case of master it was doing even three queries (wal_position two times). We will issue one SELECT for every HA loop and cache the results.	2018-01-04 11:17:43 +01:00
Ants Aasma	15d1767402	Some improvements to patronictl (#571 ) * Use scope from config file when listing members * Add version command to patronictl * Only delete leader on shutdown when we have the lock to avoid exceptions when leader key does not exist * Add a timestamp option to list command. * YAML format for patronictl output * Fix API request to get version	2018-01-04 10:35:22 +01:00
Alexander Kukushkin	0e01bb33bb	Improve patronictl reinit (#576 ) Make it possible to cancel a running task if you want to reinitialize replica. There are two possible ways to trigger it: 1. patronictl will ask whether you want to cancel already running task if an attempt to trigger reinitialize has failed 2. if you are using `--force` argument with `patronictl reinit`	2018-01-04 10:31:44 +01:00
Alexander Kukushkin	b6425cab85	Allow to specify multiple hosts for etcd (#589 ) This list will be used for initial discovery of etcd cluster members. If for some reason during work this list of hosts has been exhausted (during work), Patroni will return to initial list. In addition to that improve ipv6 compatibility by using a special function for splitting host and port. Fixes https://github.com/zalando/patroni/issues/523	2018-01-04 10:25:06 +01:00
Alexander Kukushkin	4328c15010	Make Patroni Kubernetes native (#500 ) * Use ConfigMaps or Endpoins for leader elections and to keep cluster state * Label pods with a postgres role * change behavior of pip install. From now on it will not install all dependencies, you have to specify explicitly DCS you want to use Patroni with: `pip install patroni[etcd,zookeeper,kubernetes]`	2017-12-08 16:55:00 +01:00
Alexander Kukushkin	bd847fd2cc	Patronictl extended info (#567 ) * Show information about scheduled failover and maintenance mode when showing list of cluster members. Fixes https://github.com/zalando/patroni/issues/557 * Fix postgres version check functions (postgres 10 and above compatibility) and apply pep8 formatting to the tests. * Bump some configuration parameters to match with postgres 10 defaults. * Fix name of contributor in release notes.	2017-11-28 12:10:05 +01:00
Ants Aasma	5da0e12353	Factor out postmaster process (#561 ) Introduces a PostmasterProcess object that identifies a running process via pid and start time. When pid file is parsed and the correct process identified this object is passed around. When the process goes away we try to find a new one in case somebody restarted postgres behind our back.	2017-11-23 14:36:23 +01:00
Alexander Kukushkin	a89a902f4a	Bump version and write release notes (#560 ) and implement missing unit-tests	2017-11-10 11:48:50 +01:00
Ants Aasma	7367b7c74a	Verify process start time when checking if postgres is running. (#549 ) After a crash that doesn't clean up postmaster.pid there could be a new process with the same pid resulting in a false positive for is_running(), which will lead to all kinds of bad behavior. Fixes #548	2017-11-09 15:36:05 +01:00

... 3 4 5 6 7 ...

735 Commits