Commits · ed851860b4552fc8963ecf71eab9f6f7a5c19d74 · nexedi / linux

31 May, 2014 1 commit

blk-mq: push IPI or local end_io decision to __blk_mq_complete_request() · ed851860

Jens Axboe authored May 30, 2014

We have callers outside of the blk-mq proper (like timeouts) that
want to call __blk_mq_complete_request(), so rename the function
and put the decision code for whether to use ->softirq_done_fn
or blk_mq_endio() into __blk_mq_complete_request().

This also makes the interface more logical again.
blk_mq_complete_request() attempts to atomically mark the request
completed, and calls __blk_mq_complete_request() if successful.
__blk_mq_complete_request() then just ends the request.
Signed-off-by: Jens Axboe <axboe@fb.com>

ed851860

30 May, 2014 5 commits

blk-mq: remember to start timeout handler for direct queue · feff6894

Jens Axboe authored May 30, 2014

Commit 07068d5b added a direct-to-hw-queue mode, but this mode
needs to remember to add the request timeout handler as well.
Without it, we don't track timeouts for these requests.
Signed-off-by: Jens Axboe <axboe@fb.com>

feff6894

block: ensure that the timer is always added · c7bca418

Jens Axboe authored May 30, 2014

Commit f793aa53 relaxed the timer addition a little too much.
If the timer isn't pending, we always need to add it.
Signed-off-by: Jens Axboe <axboe@fb.com>

c7bca418

blk-mq: blk_mq_unregister_hctx() can be static · ee3c5db0

Fengguang Wu authored May 30, 2014

CC: Jens Axboe <axboe@kernel.dk>
Signed-off-by: Fengguang Wu <fengguang.wu@intel.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

ee3c5db0

blk-mq: make the sysfs mq/ layout reflect current mappings · 67aec14c

Jens Axboe authored May 30, 2014

Currently blk-mq registers all the hardware queues in sysfs,
regardless of whether it uses them (e.g. they have CPU mappings)
or not. The unused hardware queues lack the cpux/ directories,
and the other sysfs entries (like active, pending, etc) are all
zeroes.

Change this so that sysfs correctly reflects the current mappings
of the hardware queues.
Signed-off-by: Jens Axboe <axboe@fb.com>

67aec14c

blk-mq: blk_mq_tag_to_rq should handle flush request · 22302375

Shaohua Li authored May 30, 2014

flush request is special, which borrows the tag from the parent
request. Hence blk_mq_tag_to_rq needs special handling to return
the flush request from the tag.
Signed-off-by: Shaohua Li <shli@fusionio.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

22302375

29 May, 2014 4 commits

block: remove dead code in scsi_ioctl:blk_verify_command · da52f22f

Dave Jones authored May 29, 2014

filter gets assigned the address of blk_default_cmd_filter on
entry to this function, so the !filter condition can never be true.
Signed-off-by: Dave Jones <davej@redhat.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

da52f22f

blk-mq: request initialization optimizations · 4b570521

Jens Axboe authored May 29, 2014

We currently clear a lot more than we need to, so make that a bit
more clever. Make some of the init dependent on features, like
only setting start_time if we are going to use it.
Signed-off-by: Jens Axboe <axboe@fb.com>

4b570521

block: add queue flag for disabling SG merging · 05f1dd53

Jens Axboe authored May 29, 2014

If devices are not SG starved, we waste a lot of time potentially
collapsing SG segments. Enough that 1.5% of the CPU time goes
to this, at only 400K IOPS. Add a queue flag, QUEUE_FLAG_NO_SG_MERGE,
which just returns the number of vectors in a bio instead of looping
over all segments and checking for collapsible ones.

Add a BLK_MQ_F_SG_MERGE flag so that drivers can opt-in on the sg
merging, if they so desire.
Signed-off-by: Jens Axboe <axboe@fb.com>

05f1dd53

block: remove 'magic' from struct blk_plug · 4d92a9be

Jens Axboe authored May 29, 2014

I don't think we've ever caught any bugs with this, and there's the
list poisoning for the plug lists to catch uninitialized cases.
So remove the magic member and save 8 bytes in the struct.
Signed-off-by: Jens Axboe <axboe@fb.com>

4d92a9be

28 May, 2014 9 commits

blk-mq: remove alloc_hctx and free_hctx methods · cdef54dd

Christoph Hellwig authored May 28, 2014

There is no need for drivers to control hardware context allocation
now that we do the context to node mapping in common code.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

cdef54dd

blk-mq: add file comments and update copyright notices · 75bb4625

Jens Axboe authored May 28, 2014

None of the blk-mq files have an explanatory comment at the top
for what that particular file does. Add that and add appropriate
copyright notices as well.
Signed-off-by: Jens Axboe <axboe@fb.com>

75bb4625

blk-mq: remove blk_mq_alloc_request_pinned · d852564f

Christoph Hellwig authored May 27, 2014

We now only have one caller left and can open code it there in a cleaner
way.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

d852564f

blk-mq: do not use blk_mq_alloc_request_pinned in blk_mq_map_request · 793597a6

Christoph Hellwig authored May 27, 2014

We already do a non-blocking allocation in blk_mq_map_request, no need
to repeat it.  Just call __blk_mq_alloc_request to wait directly.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

793597a6

blk-mq: remove blk_mq_wait_for_tags · a3bd7756

Christoph Hellwig authored May 27, 2014

The current logic for blocking tag allocation is rather confusing, as we
first allocated and then free again a tag in blk_mq_wait_for_tags, just
to attempt a non-blocking allocation and then repeat if someone else
managed to grab the tag before us.

Instead change blk_mq_alloc_request_pinned to simply do a blocking tag
allocation itself and use the request we get back from it.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

a3bd7756

blk-mq: initialize request in __blk_mq_alloc_request · 5dee8577

Christoph Hellwig authored May 27, 2014

Both callers if __blk_mq_alloc_request want to initialize the request, so
lift it into the common path.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

5dee8577

blk-mq: merge blk_mq_alloc_reserved_request into blk_mq_alloc_request · 4ce01dd1

Christoph Hellwig authored May 27, 2014

Instead of having two almost identical copies of the same code just let
the callers pass in the reserved flag directly.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

4ce01dd1

blk-mq: add helper to insert requests from irq context · 6fca6a61

Christoph Hellwig authored May 28, 2014

Both the cache flush state machine and the SCSI midlayer want to submit
requests from irq context, and the current per-request requeue_work
unfortunately causes corruption due to sharing with the csd field for
flushes.  Replace them with a per-request_queue list of requests to
be requeued.

Based on an earlier test by Ming Lei.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Reported-by: Ming Lei <tom.leiming@gmail.com>
Tested-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

6fca6a61

blk-mq: remove stale comment for blk_mq_complete_request() · 7738dac4

Jens Axboe authored May 28, 2014

It works for both IPI and local completions as of commit
95f09684.
Signed-off-by: Jens Axboe <axboe@fb.com>

7738dac4

27 May, 2014 5 commits

blk-mq: allow non-softirq completions · 95f09684

Jens Axboe authored May 27, 2014

Right now we export two ways of completing a request:

1) blk_mq_complete_request(). This uses an IPI (if needed) and
   completes through q->softirq_done_fn(). It also works with
   timeouts.

2) blk_mq_end_io(). This completes inline, and ignores any timeout
   state of the request.

Let blk_mq_complete_request() handle non-softirq_done_fn completions
as well, by just completing inline. If a driver has enough completion
ports to place completions correctly, it need not define a
mq_ops->complete() and we can avoid an indirect function call by
doing the completion inline.
Signed-off-by: Jens Axboe <axboe@fb.com>

95f09684

blk-mq: pass in suggested NUMA node to ->alloc_hctx() · f14bbe77

Jens Axboe authored May 27, 2014

Drivers currently have to figure this out on their own, and they
are missing information to do it properly. The ones that did
attempt to do it, do it wrong.

So just pass in the suggested node directly to the alloc
function.
Signed-off-by: Jens Axboe <axboe@fb.com>

f14bbe77

block: only allocate/free mq_usage_counter in blk-mq · 3d2936f4

Ming Lei authored May 27, 2014

The percpu counter is only used for blk-mq, so move
its allocation and free inside blk-mq, and don't
allocate it for legacy queue device.
Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

3d2936f4

blk-mq: avoid code duplication · 624dbe47

Ming Lei authored May 27, 2014

blk_mq_exit_hw_queues() and blk_mq_free_hw_queues()
are introduced to avoid code duplication.
Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

624dbe47

blk-mq: fix leak of hctx->ctx_map · 1f9f07e9

Ming Lei authored May 27, 2014

hctx->ctx_map should have been freed inside blk_mq_free_queue().
Signed-off-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

1f9f07e9

26 May, 2014 2 commits

block/blk-lib.c: make __blkdev_issue_zeroout static · 35086784

Fabian Frederick authored May 26, 2014

__blkdev_issue_zeroout is only used in blk-lib.c

Cc: Jens Axboe <axboe@kernel.dk>
Cc: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Fabian Frederick <fabf@skynet.be>
Signed-off-by: Jens Axboe <axboe@fb.com>

35086784

blk-mq: idle all hardware contexts before freeing a queue · 19c5d84f

Christoph Hellwig authored May 26, 2014

Without this we can leak the active_queues reference if a command is
freed while it is considered active.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <axboe@fb.com>

19c5d84f

23 May, 2014 2 commits

blk-mq: allow setting of per-request timeouts · c22d9d8a

Jens Axboe authored May 23, 2014

Currently blk-mq uses the queue timeout for all requests. But
for some commands, drivers may want to set a specific timeout
for special requests. Allow this to be passed in through
request->timeout, and use it if set.
Signed-off-by: Jens Axboe <axboe@fb.com>

c22d9d8a

blk-mq: export blk_mq_tag_busy_iter · edf866b3

Sam Bradshaw authored May 23, 2014

Export the blk-mq in-flight tag iterator for driver consumption.
This is particularly useful in exception paths or SRSI where
in-flight IOs need to be cancelled and/or reissued. The NVMe driver
conversion will use this.
Signed-off-by: Sam Bradshaw <sbradshaw@micron.com>
Signed-off-by: Matias Bjørling <m@bjorling.me>
Signed-off-by: Jens Axboe <axboe@fb.com>

edf866b3

22 May, 2014 1 commit

blk-mq: split make request handler for multi and single queue · 07068d5b

Jens Axboe authored May 22, 2014

We want slightly different behavior from them:

- On single queue devices, we currently use the per-process plug
  for deferred IO and for merging.

- On multi queue devices, we don't use the per-process plug, but
  we want to go straight to hardware for SYNC IO.

Split blk_mq_make_request() into a blk_sq_make_request() for single
queue devices, and retain blk_mq_make_request() for multi queue
devices. Then we don't need multiple checks for q->nr_hw_queues
in the request mapping.
Signed-off-by: Jens Axboe <axboe@fb.com>

07068d5b

21 May, 2014 2 commits

blk-mq: save memory by freeing requests on unused hardware queues · 484b4061

Jens Axboe authored May 21, 2014

Depending on the topology of the machine and the number of queues
exposed by a device, we can end up in a situation where some of
the hardware queues are unused (as in, they don't map to any
software queues). For this case, free up the memory used by the
request map, as we will not use it. This can be a substantial
amount of memory, depending on the number of queues vs CPUs and
the queue depth of the device.
Signed-off-by: Jens Axboe <axboe@fb.com>

484b4061

blk-mq: allow the hctx cpu hotplug notifier to return errors · e814e71b

Jens Axboe authored May 21, 2014

Prepare this for the next patch which adds more smarts in the
plugging logic, so that we can save some memory.
Signed-off-by: Jens Axboe <axboe@fb.com>

e814e71b

20 May, 2014 5 commits

blk-mq: Micro-optimize blk_queue_nomerges() check · da41a589

Robert Elliott authored May 20, 2014

In blk_mq_make_request(), do the blk_queue_nomerges() check
outside the call to blk_attempt_plug_merge() to eliminate
function call overhead when nomerges=2 (disabled)
Signed-off-by: Robert Elliott <elliott@hp.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

da41a589

blk-mq: initialize q->nr_requests after calling blk_queue_make_request() · eba71768

Jens Axboe authored May 20, 2014

blk_queue_make_requests() overwrites our set value for q->nr_requests,
turning it into the default of 128. Set this appropriately after
initializing queue values in blk_queue_make_request().
Signed-off-by: Jens Axboe <axboe@fb.com>

eba71768

blk-mq: allow changing of queue depth through sysfs · e3a2b3f9

Jens Axboe authored May 20, 2014

For request_fn based devices, the block layer exports a 'nr_requests'
file through sysfs to allow adjusting of queue depth on the fly.
Currently this returns -EINVAL for blk-mq, since it's not wired up.
Wire this up for blk-mq, so that it now also always dynamic
adjustments of the allowed queue depth for any given block device
managed by blk-mq.
Signed-off-by: Jens Axboe <axboe@fb.com>

e3a2b3f9

htmldocs: fix bio.c location · 64b14519

Jens Axboe authored May 20, 2014

Commit f9c78b2b moved bio.c from fs/ to block/, but didn't
update the docbook location. Fix that up.
Signed-off-by: Jens Axboe <axboe@fb.com>

64b14519

block: move mm/bounce.c to block/ · 719c555f

Jens Axboe authored May 19, 2014

Continue moving some of the block files that are scattered around.
bounce.c contains only code for bouncing the contents of a bio.
It's block proper code, not mm code.
Suggested-by: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: Jens Axboe <axboe@fb.com>

719c555f

19 May, 2014 4 commits

Merge branch 'for-3.16/blk-mq-tagging' into for-3.16/core · 39a9f97e
Jens Axboe authored May 19, 2014
```
Signed-off-by: Jens Axboe <axboe@fb.com>

Conflicts:
	block/blk-mq-tag.c
```
39a9f97e

blk-mq: switch ctx pending map to the sparser blk_align_bitmap · 1429d7c9

Jens Axboe authored May 19, 2014

Each hardware queue has a bitmap of software queues with pending
requests. When new IO is queued on a software queue, the bit is
set, and when IO is pruned on a hardware queue run, the bit is
cleared. This causes a lot of traffic. Switch this from the regular
BITS_PER_LONG bitmap to a sparser layout, similarly to what was
done for blk-mq tagging.

20% performance increase was observed for single threaded IO, and
about 15% performanc increase on multiple threads driving the
same device.
Signed-off-by: Jens Axboe <axboe@fb.com>

1429d7c9

blk-mq: move the cache friendly bitmap type of out blk-mq-tag · e93ecf60
Jens Axboe authored May 19, 2014
```
We will use it for the pending list in blk-mq core as well.
Signed-off-by: Jens Axboe <axboe@fb.com>
```
e93ecf60

block: move ioprio.c from fs/ to block/ · 2667bcbb

Jens Axboe authored May 19, 2014

Like commit f9c78b2b, move this block related file outside
of fs/ and into the core block directory, block/.
Signed-off-by: Jens Axboe <axboe@fb.com>

2667bcbb