Commits · 0d2602ca30e410e84e8bdf05c84ed5688e0a5a44 · Mirrors / git.kernel.org / pub_scm_linux_kernel_git_torvalds_linux

May 14, 2014

blk-mq: improve support for shared tags maps · 0d2602ca

Jens Axboe authored May 13, 2014



This adds support for active queue tracking, meaning that the
blk-mq tagging maintains a count of active users of a tag set.
This allows us to maintain a notion of fairness between users,
so that we can distribute the tag depth evenly without starving
some users while allowing others to try unfair deep queues.

If sharing of a tag set is detected, each hardware queue will
track the depth of its own queue. And if this exceeds the total
depth divided by the number of active queues, the user is actively
throttled down.

The active queue count is done lazily to avoid bouncing that data
between submitter and completer. Each hardware queue gets marked
active when it allocates its first tag, and gets marked inactive
when 1) the last tag is cleared, and 2) the queue timeout grace
period has passed.

Signed-off-by: Jens Axboe <axboe@fb.com>

0d2602ca

May 11, 2014

blk-mq: bitmap tag: cleanup blk_mq_init_tags · 1f236ab2

Ming Lei authored May 11, 2014



Both nr_cache and nr_tags arn't needed for bitmap tag anymore.

Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

1f236ab2

blk-mq: bitmap tag: select random tag betweet 0 and (depth - 1) · 9d3d21ae

Ming Lei authored May 10, 2014



The selected tag should be selected at random between 0 and
(depth - 1) with probability 1/depth, instead between 0 and
(depth - 2) with probability 1/(depth - 1).

Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

9d3d21ae

blk-mq: bitmap tag: remove barrier in bt_clear_tag() · 60f2df8a

Ming Lei authored May 11, 2014



The barrier isn't necessary because both atomic_dec_and_test()
and wake_up() implicate one barrier.

Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

60f2df8a

blk-mq: bitmap tag: use clear_bit_unlock in bt_clear_tag() · 0289b2e1

Ming Lei authored May 11, 2014



The unlock memory barrier need to order access to req in free
path and clearing tag bit, otherwise either request free path
may see a allocated request, or initialized request in allocate
path might be modified by the ongoing free path.

Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

0289b2e1

May 10, 2014

blk-mq: use sparser tag layout for lower queue depth · 59d13bf5

Jens Axboe authored May 09, 2014



For best performance, spreading tags over multiple cachelines
makes the tagging more efficient on multicore systems. But since
we have 8 * sizeof(unsigned long) tags per cacheline, we don't
always get a nice spread.

Attempt to spread the tags over at least 4 cachelines, using fewer
number of bits per unsigned long if we have to. This improves
tagging performance in setups with 32-128 tags. For higher depths,
the spread is the same as before (BITS_PER_LONG tags per cacheline).

Signed-off-by: Jens Axboe <axboe@fb.com>

59d13bf5

May 09, 2014

blk-mq: implement new and more efficient tagging scheme · 4bb659b1

Jens Axboe authored May 09, 2014



blk-mq currently uses percpu_ida for tag allocation. But that only
works well if the ratio between tag space and number of CPUs is
sufficiently high. For most devices and systems, that is not the
case. The end result if that we either only utilize the tag space
partially, or we end up attempting to fully exhaust it and run
into lots of lock contention with stealing between CPUs. This is
not optimal.

This new tagging scheme is a hybrid bitmap allocator. It uses
two tricks to both be SMP friendly and allow full exhaustion
of the space:

1) We cache the last allocated (or freed) tag on a per blk-mq
   software context basis. This allows us to limit the space
   we have to search. The key element here is not caching it
   in the shared tag structure, otherwise we end up dirtying
   more shared cache lines on each allocate/free operation.

2) The tag space is split into cache line sized groups, and
   each context will start off randomly in that space. Even up
   to full utilization of the space, this divides the tag users
   efficiently into cache line groups, avoiding dirtying the same
   one both between allocators and between allocator and freeer.

This scheme shows drastically better behaviour, both on small
tag spaces but on large ones as well. It has been tested extensively
to show better performance for all the cases blk-mq cares about.

Signed-off-by: Jens Axboe <axboe@fb.com>

4bb659b1

blk-mq: initialize struct request fields individually · af76e555

Christoph Hellwig authored May 06, 2014



This allows us to avoid a non-atomic memset over ->atomic_flags as well
as killing lots of duplicate initializations.

Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

af76e555

blk-mq: update a hotplug comment for grammar · 9fccfed8
Jens Axboe authored May 08, 2014
```
Signed-off-by: Jens Axboe <axboe@fb.com>
```
9fccfed8

May 08, 2014

blk-mq: add basic round-robin of what CPU to queue workqueue work on · 506e931f

Jens Axboe authored May 07, 2014



Right now we just pick the first CPU in the mask, but that can
easily overload that one. Add some basic batching and round-robin
all the entries in the mask instead.

Signed-off-by: Jens Axboe <axboe@fb.com>

506e931f

May 03, 2014

block/blk-throttle.c: fix return of 0/1 with return type bool · 5cf8c227

Fabian Frederick authored May 02, 2014



Fix 4 coccinelle warnings.

Cc: Jens Axboe <axboe@kernel.dk>
Cc: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Fabian Frederick <fabf@skynet.be>
Signed-off-by: Jens Axboe <axboe@fb.com>

5cf8c227

block/blk-iopoll.c: use iop instead of iopoll · 5214e33c

Fabian Frederick authored May 02, 2014



All blk_iopoll functions use iop for parent iopoll structure except
blk_iopoll_complete.This also fixes one kernel-doc warning.

Cc: Jens Axboe <axboe@kernel.dk>
Cc: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Fabian Frederick <fabf@skynet.be>
Signed-off-by: Jens Axboe <axboe@fb.com>

5214e33c

blk-mq: remove extra requeue trace · 74814b1c

Jens Axboe authored May 02, 2014



We already issue a blktrace requeue event in
__blk_mq_requeue_request(), don't do it from the original caller
as well.

Signed-off-by: Jens Axboe <axboe@fb.com>

74814b1c

May 01, 2014

block: Fix format string mismatch in cfq-iosched.c · 176167ad

Masanari Iida authored Apr 28, 2014



Fix format string mismatch in cfq_var_show()

Signed-off-by: Masanari Iida <standby24x7@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

176167ad

blk-mq: refactor request insertion/merging · c6d600c6

Jens Axboe authored Apr 30, 2014



Refactor the logic around adding a new bio to a software queue,
so we nest the ctx->lock where we really need it (merge and
insertion) and don't hold it when we don't (init and IO start
accounting).

Signed-off-by: Jens Axboe <axboe@fb.com>

c6d600c6

blk-mq remove debug BUG_ON() when draining software queues · 98bc1f27
Jens Axboe authored Apr 30, 2014
```
It's never been of any use, lets get rid of it.

Signed-off-by: Jens Axboe <axboe@fb.com>
```
98bc1f27

Apr 30, 2014

blk-mq: fix waiting for reserved tags · 5810d903

Jens Axboe authored Apr 29, 2014



blk_mq_wait_for_tags() is only able to wait for "normal" tags,
not reserved tags. Pass in which one we should attempt to get
a tag for, so that waiting for reserved tags will work.

Reserved tags are used for internal commands, which are usually
serialized. Hence no waiting generally takes place, but we should
ensure that it actually works if users need that functionality.

Signed-off-by: Jens Axboe <axboe@fb.com>

5810d903

Apr 28, 2014

random: export add_disk_randomness · bdcfa3e5

Christoph Hellwig authored Apr 25, 2014



This will be needed for pending changes to the scsi midlayer that now
calls lower level block APIs, as well as any blk-mq driver that wants to
contribute to the random pool.

Signed-off-by: Christoph Hellwig <hch@lst.de>
Acked-by: "Theodore Ts'o" <tytso@mit.edu>
Signed-off-by: Jens Axboe <axboe@fb.com>

bdcfa3e5

Apr 25, 2014

block: fold __blk_add_timer into blk_add_timer · c4a634f4
Christoph Hellwig authored Apr 25, 2014
```
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>
```
c4a634f4

blk-mq: respect rq_affinity · 38535201

Christoph Hellwig authored Apr 25, 2014



The blk-mq code is using it's own version of the I/O completion affinity
tunables, which causes a few issues:

 - the rq_affinity sysfs file doesn't work for blk-mq devices, even if it
   still is present, thus breaking existing tuning setups.
 - the rq_affinity = 1 mode, which is the defauly for legacy request based
   drivers isn't implemented at all.
 - blk-mq drivers don't implement any completion affinity with the default
   flag settings.

This patches removes the blk-mq ipi_redirect flag and sysfs file, as well
as the internal BLK_MQ_F_SHOULD_IPI flag and replaces it with code that
respects the queue-wide rq_affinity flags and also implements the
rq_affinity = 1 mode.

This means I/O completion affinity can now only be tuned block-queue wide
instead of per context, which seems more sensible to me anyway.

Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

38535201

Apr 24, 2014

blk-mq: fix race with timeouts and requeue events · 87ee7b11

Jens Axboe authored Apr 24, 2014



If a requeue event races with a timeout, we can get into the
situation where we attempt to complete a request from the
timeout handler when it's not start anymore. This causes a crash.
So have the timeout handler check that REQ_ATOM_STARTED is still
set on the request - if not, we ignore the event. If this happens,
the request has now been marked as complete. As a consequence, we
need to ensure to clear REQ_ATOM_COMPLETE in blk_mq_start_request(),
as to maintain proper request state.

Signed-off-by: Jens Axboe <axboe@fb.com>

87ee7b11

Revert "blk-mq: initialize req->q in allocation" · 70ab0b2d

Jens Axboe authored Apr 24, 2014

This reverts commit 6a3c8a3a.

We need selective clearing of the request to make the init-at-free
time completely safe. Otherwise we end up stomping on
rq->atomic_flags, which we don't want to do.

70ab0b2d

blk-mq: fix leak of set->tags · 981bd189

Ming Lei authored Apr 24, 2014



set->tags should be freed in blk_mq_free_tag_set().

Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

981bd189

Apr 23, 2014

fs/bio.c: remove nr_segs (unused function parameter) · 7410b3c6

Fabian Frederick authored Apr 22, 2014

nr_segs is no longer used in bio_alloc_map_data since c8db4448


("block: Don't save/copy bvec array anymore")

Signed-off-by: Fabian Frederick <fabf@skynet.be>
Cc: Jens Axboe <axboe@kernel.dk>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Jens Axboe <axboe@fb.com>

7410b3c6

fs/bio: remove bs paramater in biovec_create_pool · a6c39cb4

Fabian Frederick authored Apr 22, 2014

bs is no longer used in biovec_create_pool since 9f060e22

 ("block:
Convert integrity to bvec_alloc_bs()")

Signed-off-by: Fabian Frederick <fabf@skynet.be>
Cc: Jens Axboe <axboe@kernel.dk>
Signed-off-by: Jens Axboe <axboe@fb.com>

a6c39cb4

Apr 22, 2014

block/blk-throttle.c: add static to blk_throtl_dispatch_work_fn · 8876e140

Fabian Frederick authored Apr 17, 2014



blk_throtl_dispatch_work_fn is only used in blk-throttle.c

Cc: Jens Axboe <axboe@kernel.dk>
Cc: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Fabian Frederick <fabf@skynet.be>
Signed-off-by: Jens Axboe <axboe@fb.com>

8876e140

fs: fix new kernel-doc warnings in fs/bio.c · 1051a902

Randy Dunlap authored Apr 20, 2014



Fix new kernel-doc warnings in fs/bio.c:

Warning(fs/bio.c:316): No description found for parameter 'bio'
Warning(fs/bio.c:316): No description found for parameter 'parent'

Signed-off-by: Randy Dunlap <rdunlap@infradead.org>
Signed-off-by: Jens Axboe <axboe@fb.com>

1051a902

blk-mq: initialize req->q in allocation · 6a3c8a3a

Ming Lei authored Apr 19, 2014



The patch basically reverts the patch of(blk-mq:
initialize request on allocation) in Jens's tree(already
in -next), and only initialize req->q in allocation
for two reasons:

	- presumed cache hotness on completion
	- blk_rq_tagged(rq) depends on reset of req->mq_ctx

Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

6a3c8a3a

blk-mq: user (1 << order) to implement order_to_size() · 4ca08500

Ming Lei authored Apr 19, 2014



Cc: Jörg-Volker Peetz <jvpeetz@web.de>
Cc: Max Filippov <jcmvbkbc@gmail.com>
Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

4ca08500

blk-mq: fix allocation of set->tags · 48479005

Ming Lei authored Apr 19, 2014



type of set->tags is struct blk_mq_tags **.

Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

48479005

blk-mq: free hctx->ctx_map when init failed · 11471e0d

Ming Lei authored Apr 19, 2014



Avoid memory leak in the failure path.

Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

11471e0d

Apr 17, 2014

sd/skd: stuff discard page in request->completion_data · dc4a9307

Jens Axboe authored Apr 16, 2014



Store the pointer to the page there, so we can always safely
reference it from end_io context where ->bio may have been
cleared.

Signed-off-by: Jens Axboe <axboe@fb.com>

dc4a9307

jsflash: missed conversion from rq->buffer to bio_data(rq->bio) · fb1be433
Jens Axboe authored Apr 16, 2014
```
Signed-off-by: Jens Axboe <axboe@fb.com>
```
fb1be433

block: relax when to modify the timeout timer · f793aa53

Jens Axboe authored Apr 16, 2014



Since we are now, by default, applying timer slack to expiry times,
the logic for when to modify a timer in the block code is suboptimal.
The block layer keeps a forward rolling timer per queue for all
requests, and modifies this timer if a request has a shorter timeout
than what the current expiry time is. However, this breaks down
when our rounded timer values get applied slack. Then each new
request ends up modifying the timer, since we're still a little
in front of the timer + slack.

Fix this by allowing a tolerance of HZ / 2, the timeout handling
doesn't need to be very precise. This drastically cuts down
the number of timer modifications we have to make.

Signed-off-by: Jens Axboe <axboe@fb.com>

f793aa53

block: export blk_finish_request · 12120077

Christoph Hellwig authored Apr 16, 2014



This allows to mirror the blk-mq code flow for more a more readable I/O
completion handler in SCSI.

Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

12120077

blk-mq: rename mq_flush_work struct request member · f88a164b

Christoph Hellwig authored Apr 16, 2014



We will use this work_struct to requeue scsi commands from the
completion handler as well, so give it a more generic name.

Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

f88a164b

blk-mq: add blk_mq_requeue_request · ed0791b2

Christoph Hellwig authored Apr 16, 2014



This allows to requeue a request that has been accepted by ->queue_rq
earlier.  This is needed by the SCSI layer in various error conditions.

The existing internal blk_mq_requeue_request is renamed to
__blk_mq_requeue_request as it is a lower level building block for this
funtionality.

Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

ed0791b2

blk-mq: add blk_mq_start_hw_queues · 2f268556

Christoph Hellwig authored Apr 16, 2014



Add a helper to unconditionally kick contexts of a queue.  This will
be needed by the SCSI layer to provide fair queueing between multiple
devices on a single host.

Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

2f268556

blk-mq: add blk_mq_delay_queue · 70f4db63

Christoph Hellwig authored Apr 16, 2014



Add a blk-mq equivalent to blk_delay_queue so that the scsi layer can ask
to be kicked again after a delay.

Signed-off-by: Christoph Hellwig <hch@lst.de>

Modified by me to kill the unnecessary preempt disable/enable
in the delayed workqueue handler.

Signed-off-by: Jens Axboe <axboe@fb.com>

70f4db63

blk-mq: add async parameter to blk_mq_start_stopped_hw_queues · 1b4a3258
Christoph Hellwig authored Apr 16, 2014
```
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>
```
1b4a3258