freebsd-nq

Author	SHA1	Message	Date
Alexander Motin	40ea77a036	Merge GEOM direct dispatch changes from the projects/camlock branch. When safety requirements are met, it allows to avoid passing I/O requests to GEOM g_up/g_down thread, executing them directly in the caller context. That allows to avoid CPU bottlenecks in g_up/g_down threads, plus avoid several context switches per I/O. The defined now safety requirements are: - caller should not hold any locks and should be reenterable; - callee should not depend on GEOM dual-threaded concurency semantics; - on the way down, if request is unmapped while callee doesn't support it, the context should be sleepable; - kernel thread stack usage should be below 50%. To keep compatibility with GEOM classes not meeting above requirements new provider and consumer flags added: - G_CF_DIRECT_SEND -- consumer code meets caller requirements (request); - G_CF_DIRECT_RECEIVE -- consumer code meets callee requirements (done); - G_PF_DIRECT_SEND -- provider code meets caller requirements (done); - G_PF_DIRECT_RECEIVE -- provider code meets callee requirements (request). Capable GEOM class can set them, allowing direct dispatch in cases where it is safe. If any of requirements are not met, request is queued to g_up or g_down thread same as before. Such GEOM classes were reviewed and updated to support direct dispatch: CONCAT, DEV, DISK, GATE, MD, MIRROR, MULTIPATH, NOP, PART, RAID, STRIPE, VFS, ZERO, ZFS::VDEV, ZFS::ZVOL, all classes based on g_slice KPI (LABEL, MAP, FLASHMAP, etc). To declare direct completion capability disk(9) KPI got new flag equivalent to G_PF_DIRECT_SEND -- DISKFLAG_DIRECT_COMPLETION. da(4) and ada(4) disk drivers got it set now thanks to earlier CAM locking work. This change more then twice increases peak block storage performance on systems with manu CPUs, together with earlier CAM locking changes reaching more then 1 million IOPS (512 byte raw reads from 16 SATA SSDs on 4 HBAs to 256 user-level threads). Sponsored by: iXsystems, Inc. MFC after: 2 months	2013-10-22 08:22:19 +00:00
Alexander Motin	a93c0ed463	Remove extra bio_data and bio_length copying to child request after calling g_clone_bio(), that already copied them.	2013-03-26 05:42:12 +00:00
Pawel Jakub Dawidek	c4d2d401f8	We don't need buffer to handle BIO_DELETE, so don't check buffer size for it. This fixes handling BIO_DELETE larger than MAXPHYS.	2013-03-14 23:07:01 +00:00
Mikolaj Golub	1d9db37c77	In g_gate_dumpconf() always check the result of g_gate_hold(). This fixes "Negative sc_ref" panic possible when sysctl_kern_geom_confxml() is run simultaneously with destroying GATE device. Reviewed by: pjd MFC after: 3 days	2012-08-07 18:50:33 +00:00
Mikolaj Golub	a277f47bd2	Reorder things in g_gate_create() so at the moment when g_new_geomf() is called name is properly initialized. Discussed with: pjd MFC after: 2 weeks	2012-07-28 16:30:50 +00:00
Pawel Jakub Dawidek	e08ec03778	Extend GEOM Gate class to handle read I/O requests directly within the kernel. This will allow HAST to read directly from the local component without even communicating userland daemon. Sponsored by: Panzura, http://www.panzura.com MFC after: 1 month	2012-07-04 20:16:28 +00:00
Ed Schouten	6472ac3d8a	Mark all SYSCTL_NODEs static that have no corresponding SYSCTL_DECLs. The SYSCTL_NODE macro defines a list that stores all child-elements of that node. If there's no SYSCTL_DECL macro anywhere else, there's no reason why it shouldn't be static.	2011-11-07 15:43:11 +00:00
Andrey V. Elsukov	5d807a0e1a	Include sys/sbuf.h directly. Reviewed by: pjd	2011-07-11 05:22:31 +00:00
Pawel Jakub Dawidek	204a4e196a	Recognize BIO_FLUSH requests and pass them to userland. MFC after: 1 week	2011-05-23 21:00:37 +00:00
Pawel Jakub Dawidek	63a6c5c12b	GEOM has an internal mechanism to deal with ENOMEM errors returned via g_io_deliver(). In such case it increases 'pace' counter on each ENOMEM and reschedules the request. The 'pace' counter is decreased for each request going down, but until 'pace' is greater than zero, GEOM will handle at most 10 requests per second. For GEOM GATE users that are proxy to local GEOM providers (like ggatel(8) and HAST) we can end up with almost permanent slow down of GEOM down queue. This is because once we reach GEOM GATE queue limit, we return ENOMEM to the GEOM. This means that we have, eg. 1024 I/O requests in the GEOM GATE queue. To make room in the queue and stop returning ENOMEM we need to proceed the requests of course, but those requests are handled by userland daemons that handle them by reading/writing also from/to local GEOM providers. For example with HAST, a new requests comes to /dev/hast/data, which is GEOM GATE provider. GEOM GATE passes the request to hastd(8) and hastd(8) reads/writes from/to /dev/da0. Once we reach GEOM GATE queue limit, to free up a slot in GEOM GATE queue, hastd(8) has to read/write from/to /dev/da0, but this request will also be very slow, because GEOM now slows down all the requests. We end up with full queue that we can unload at the speed of 10 requests per second. This simply looks like a deadlock. Fix it by allowing userland daemons that work with both GEOM GATE and local GEOM providers to specify unlimited queue size, so GEOM GATE will never return ENOMEM to the GEOM. MFC after: 1 week	2011-04-02 06:56:06 +00:00
Mikolaj Golub	bd119384c7	Increase debug level on g_gate device destruction and add message on device creation. Suggested by: danger Approved by: pjd (mentor) MFC after: 3 days	2011-03-30 21:40:14 +00:00
Mikolaj Golub	baf63f65ae	In g_gate_create() there is a window between when g_gate_softc is registered in g_gate_units array and when its sc_provider field is filled. If during this period g_gate_units is accessed by another thread that is checking for provider name collision the crash is possible. Fix this by adding sc_name field to struct g_gate_softc. In g_gate_create() when g_gate_softc is created but sc_provider is still not sc_name points to provider name stored in the local array. Approved by: pjd (mentor) Reported by: Freddie Cash <fjwcash@gmail.com> MFC after: 1 week	2011-03-27 19:56:55 +00:00
Alexander Leidinger	cb08c2cc83	Add some FEATURE macros for various GEOM classes. No FreeBSD version bump, the userland application to query the features will be committed last and can serve as an indication of the availablility if needed. Sponsored by: Google Summer of Code 2010 Submitted by: kibab Reviewed by: silence on geom@ during 2 weeks X-MFC after: to be determined in last commit with code from this project	2011-02-25 10:24:35 +00:00
Pawel Jakub Dawidek	2aa15ffdab	'unit' can be negative, so use signed type for it. Found by: Coverity Prevent CID: 3731 MFC after: 3 days	2010-06-14 21:58:55 +00:00
Pawel Jakub Dawidek	15725379d0	BIO_DELETE contains range we want to delete and doesn't provide any useful data, so there is no need to copy it to userland. MFC after: 3 days	2010-06-14 21:56:24 +00:00
Pawel Jakub Dawidek	b0990a1dae	Simplify loops.	2010-03-18 13:11:43 +00:00
Pawel Jakub Dawidek	32115b105a	Please welcome HAST - Highly Avalable Storage. HAST allows to transparently store data on two physically separated machines connected over the TCP/IP network. HAST works in Primary-Secondary (Master-Backup, Master-Slave) configuration, which means that only one of the cluster nodes can be active at any given time. Only Primary node is able to handle I/O requests to HAST-managed devices. Currently HAST is limited to two cluster nodes in total. HAST operates on block level - it provides disk-like devices in /dev/hast/ directory for use by file systems and/or applications. Working on block level makes it transparent for file systems and applications. There in no difference between using HAST-provided device and raw disk, partition, etc. All of them are just regular GEOM providers in FreeBSD. For more information please consult hastd(8), hastctl(8) and hast.conf(5) manual pages, as well as http://wiki.FreeBSD.org/HAST. Sponsored by: FreeBSD Foundation Sponsored by: OMCnet Internet Service GmbH Sponsored by: TransIP BV	2010-02-18 23:16:19 +00:00
Antoine Brodin	13e403fdea	(S)LIST_HEAD_INITIALIZER takes a (S)LIST_HEAD as an argument. Fix some wrong usages. Note: this does not affect generated binaries as this argument is not used. PR: 137213 Submitted by: Eygene Ryabinkin (initial version) MFC after: 1 month	2009-12-28 22:56:30 +00:00
Pawel Jakub Dawidek	fc024f7a45	Bump copyright year.	2006-09-08 10:20:44 +00:00
Pawel Jakub Dawidek	c076790223	Use __FBSDID in .c files.	2006-09-08 10:19:24 +00:00
Pawel Jakub Dawidek	7ffb6e0f6a	Fix problems with destroy and forcible destroy functionality: - hold/release device in start/done routines, this will probably slow down things a bit, but previous code was racy; - only release device if g_gate_destroy() failed - if it succeeded device is dead and there is nothing to release; - various other changes which makes forcible destruction reliable. MFC after: 3 days	2006-09-05 21:56:00 +00:00
Pawel Jakub Dawidek	38ea96ac99	Remove trailing spaces.	2006-02-01 12:06:01 +00:00
Robert Watson	5bb84bc84b	Normalize a significant number of kernel malloc type names: - Prefer '_' to ' ', as it results in more easily parsed results in memory monitoring tools such as vmstat. - Remove punctuation that is incompatible with using memory type names as file names, such as '/' characters. - Disambiguate some collisions by adding subsystem prefixes to some memory types. - Generally prefer lower case to upper case. - If the same type is defined in multiple architecture directories, attempt to use the same name in additional cases. Not all instances were caught in this change, so more work is required to finish this conversion. Similar changes are required for UMA zone names.	2005-10-31 15:41:29 +00:00
Pawel Jakub Dawidek	84436f14c4	Add CANCEL command which allows to remove one request from the queue or all requests from the queue if request number is not given. Bump version number. Approved by: re (scottl)	2005-07-08 21:08:53 +00:00
Pawel Jakub Dawidek	0218292cdf	Update copyright in files changed this year.	2005-02-16 22:14:52 +00:00
Pawel Jakub Dawidek	ccbef85dd0	Remove mutex asserion from g_gate_find(). We don't want g_gate_list_mtx mutex to be held here, because we want speed here.	2005-02-16 16:13:56 +00:00
Pawel Jakub Dawidek	f906581296	Remove TDP_GEOM flag from thread after ggate device creation. This flag means "wait for all pending requests before returning to userland". There are pending events for sure, because we just created new provider and other classes want to taste it, but we cannot answer on I/O requests until we're here.	2005-02-16 16:12:28 +00:00
Pawel Jakub Dawidek	35f855d9f9	Fix typo. We want to unlock mutex here. Submitted by: Andreas Kohn <andreas.kohn@gmail.com> MFC after: 1 week	2005-02-12 16:19:03 +00:00
Pawel Jakub Dawidek	e35d3a7828	- Remove g_gate_hold()/g_gate_release() from start/done paths. It saves 4 mutex operations per I/O requests. - Use only one mutex to protect both (incoming and outgoing) queue. As MUTEX_PROFILING(9) shows, there is no big contention for this lock. - Protect sc_queue_count with queue mutex, instead of doing atomic operations on it. - Remove DROP_GIANT()/PICKUP_GIANT() - ggate is marked as MPSAFE and no Giant there.	2005-02-09 08:29:39 +00:00
Pawel Jakub Dawidek	662a4e5878	- Use bioq_insert_tail()/bioq_insert_head() instead of bioq_disksort(). - Improve mediasize checking. MFC after: 1 week	2005-02-05 00:30:08 +00:00
Pawel Jakub Dawidek	a17dd95f14	- Add missing Giant drop before acquiring the topology lock. - Move DROP_GIANT()/PICKUP_GIANT() to g_gate_ioctl().	2004-11-23 11:18:26 +00:00
Pawel Jakub Dawidek	c7e17f4bbe	Unlock g_gate_list_mtx mutex when we cannot allocate unit number. MT5 candidate. PR: kern/72253 Submitted by: Ivan Voras <ivoras@fer.hr>	2004-10-02 15:03:26 +00:00
Poul-Henning Kamp	5721c9c76a	Tag all geom classes in the tree with a version number.	2004-08-08 07:57:53 +00:00
Poul-Henning Kamp	3e019deaed	Do a pass over all modules in the kernel and make them return EOPNOTSUPP for unknown events. A number of modules return EINVAL in this instance, and I have left those alone for now and instead taught MOD_QUIESCE to accept this as "didn't do anything".	2004-07-15 08:26:07 +00:00
Pawel Jakub Dawidek	9ecdd506ea	Remove unused argument for good.	2004-07-01 15:42:03 +00:00
Pawel Jakub Dawidek	55336b83e0	Introduce a hack that will make geom_gate to work with read-only mounts. Now, when trying to mount file system in read-only mode it tries to opened a device for writting to be able to update to read-write mode latter. Ehh. Discussed with: phk	2004-06-27 12:56:11 +00:00
Pawel Jakub Dawidek	47f44cb708	Don't hold topology lock while calling g_gate_release(). Found by: KASSERT()	2004-06-21 09:12:08 +00:00
Poul-Henning Kamp	89c9c53da0	Do the dreaded s/dev_t/struct cdev */ Bump __FreeBSD_version accordingly.	2004-06-16 09:47:26 +00:00
Pawel Jakub Dawidek	053271038e	Close some small wakeup<->msleep races.	2004-05-05 12:30:41 +00:00
Pawel Jakub Dawidek	b62093b274	Turn off debugging by default.	2004-05-03 21:11:54 +00:00
Pawel Jakub Dawidek	37c9eaae29	Prefer signed type over unsigned to be able to assert negative reference count.	2004-05-03 21:02:02 +00:00
Pawel Jakub Dawidek	4d1e1bf3f5	- Hold g_gate_list_mtx lock while generating/checking unit number. Found by: mtx_assert() g_gate.c:273 - Set command before returning to userland with ENOMEM error value. Found by: assert() ggatel.c:108	2004-05-03 18:06:24 +00:00
Pawel Jakub Dawidek	0d785336d1	Make it compile on 64-bit architectures. The biggest issue was that 16-bit atomic operations aren't supported on all architectures.	2004-05-02 17:57:49 +00:00
Pawel Jakub Dawidek	fe27e77251	Kernel bits of GEOM Gate.	2004-04-30 16:08:12 +00:00

44 Commits